{
  "id": 56283,
  "title": "Solution #6 overview",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56283",
  "author_name": "CPMP",
  "post_date": "2018-05-08T08:30:47.227000",
  "votes": 240,
  "comment_count": 127,
  "views": 0,
  "content": "<p>First of all, all those who managed to get decent submissions out of this huge dataset deserve kudos.  Even if you score is below Kirk's shared kernel.  What you did is way more valuable, and lessons learned here will help you later.  </p>\n\n<p>Second, sharing is great when done in good faith, and lots of people did share a lot here, too many of them to name them.  Eve people who started here, like @Samrat, shared a lot.  This is what makes this community so valuable.</p>\n\n<p>Third, thanks to <a href=\"/inversion\">@inversion</a>, Kaggle, and Talking data for organizing a very challenging, ans almost leak free competition.  I say almost because it is clear now that the test data was sorted by click time then target value.  Exploiting this was the final twist that helped some of us fare better.  But the impact is not that large, I estimate it to be about 0.0004 for me.  And it was <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\">disclosed soon enough</a> for everyone to react to it.  Thanks to <a href=\"/plantsgo\">@plantsgo</a> for sharing it soon enough.</p>\n\n<p>I didn't decide to go solo from the start, but as time went by, I saw I was making progress every day, and decided to go solo till the end.  In retrospect I am not sure it was wise, I didn't sleep much in the last week ;)  I did receive some invites to merge during the last week before deadline, and I thank people for them.  I did miss a very late invite from a top 10 team as I was away that evening.  I wonder what would have been our score if we had teamed.</p>\n\n<p>Anyway here is my solution.  Given I was solo, and given the size of the data set, which meant hours to produce a submission, I decided to focus. I focused on a single type of model, LightGBM.  I split my time roughly as follows:</p>\n\n<ul>\n<li>80% feature engineering</li>\n<li>10% making local validation as fast as possible</li>\n<li>5% hyper parameter tuning</li>\n<li>5% ensembling</li>\n</ul>\n\n<p>I spent most of my time doing feature selection, as my machine was not usable with 50 features or more.  I wish I had used the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\">trick shared by Kruegger</a>, maybe I would have been able to add more features.  However, being forced to be selective about features probably led to better models in the end.  And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by <a href=\"/authman\">@authman</a>.</p>\n\n<p>My private LB score comes from a single lgb run with 48 features that scored 0.9825 public and 0.9835 private.  I submitted a blend of this with 5 other similar models that yield 0.9828 public and 0.9837 private, but for some reason <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56234\">that sub isn't taken into account</a>.  This is OK as my rank would not change with it.  I didn't submit all the 5 models individually during the competition, but did it after to get their score.  The best single lgb run scores 0.9827 public and 0.9836 private, with 48 features.</p>\n\n<p>I mostly used a 20 core Xeon at 2.3 MHz with 64 GB RAM and 64GB swap.  That machine is a bit slow, but it scales when using 20 threads.  I also used another machine with a 4 core i7 and 2 GPU to run Keras (see below).</p>\n\n<p><strong>Validation</strong></p>\n\n<p>Let's us look at validation.  This is key.  If you don't have a good validation scheme then you rely solely on LB probing, which can easily lead to overfit.  I ended up settling on:</p>\n\n<ul>\n<li>training on day &lt;=  8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.</li>\n<li>retraining on all data using 1.2 times the number of trees found by early stopping in validation</li>\n</ul>\n\n<p>Using two validation sets was to make sure I was not overfiting to one of them.  The hours were selected to match the public and private test data.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot. I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.</p>\n\n<p>I also used LB, i.e. only kept a feature if local auc and LB improved. Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.</p>\n\n<p>This was a very effective scheme, with the hour 4 score being the same as LB score with a std difference around 0.0001.  However, it is very time consuming because of the computation of the auc metric for early stopping.  In order to speed it for feature evaluation I used two lighter ways.  First, using only day 9 data, with 5% of hour 4 data for validation, the rest for training.  This could run in less than one hour, and was used as a filter.  Only features that improved on that went to the next stage which was train on day 8 and validate on day 9, both hour 4 and other test hours.  This was also a very effective scheme, with very good correlation with LB score, but it was not effective when evaluating lag features.  I therefore switched to training on day &lt;= 8 later on.</p>\n\n<p>Another way of speeding feature evaluation was to share each feature in a separate feather file.  This way, testing a feature set only requires assembling a set of files into one dataset.  Features were mostly tested by adding them one by one, and keeping them if local validation score improved by at least 0.00005.  I also added several of them at once, then removed them one by one to see if validation score decreased.  I basically did feature selection full time for the competition, preparing experiments to be run while I was away during day, or while I was sleeping.  The machine never stopped.</p>\n\n<p><strong>Feature Engineering</strong></p>\n\n<p>Features were computed on the concatenation of train and test_supplement, sorted by click time then by original order.  Now I am not sure the second item was useful.</p>\n\n<p>I used several families of features.</p>\n\n<ul>\n<li>Only  app, and os from the original features were kept. They were handled as categorical, and were my strongest 2 features with a third category made of the hour in the day.</li>\n<li>China days.  Introduced 24 periods that start at 4 pm.  These were\nused for lag features based on previous day(s) data.</li>\n<li>User: ip, device, os triplets.  </li>\n<li>Aggregates on various feature groups, similar to what was shared in many public kernels.  Aggregates I used were count, count of unique values, delta with previous value, delta with next value.  Time to next click when grouped by user was important.  Other useful ones I didn't see in kernels: delta with previous app.</li>\n<li>Lag features, based on previous China days values.  Previous count by some grouping, and previous target mean by some grouping.  The latter was a weighted average with the overall target mean, the weights being such that groups with few rows in it had a value closer to the overall average.  This is a standard normalization in target encoding.</li>\n<li>Ratios like number of clicks per ip, app to number of click per app.</li>\n<li>Not last.  This was to capture the leak.  It is one except for rows that are not the last of their group when grouped by user, app, and click time.  I ignored channel as I think that clicks are attributed to the most recent click having same user and app as the download.</li>\n<li>Target.  This is to also capture the leak.  I modified the target in train data by sorting is_attributed within group by user, app, and click time. The combination of both ways to capture the leak led to a boost between 0.0004 and 0.0005.</li>\n<li>Matrix factorization.  This was to capture the similarity between users and app.  I use several of them.  They all start with the same approach; construct a matrix with log of click counts. I used: ip x app, user x app, and os x device x app.  These matrices are extremely sparse (most values are 0).  For the first two I used truncated svd from sklearn, which gives me latent vectors (embeddings) for ip and user.  For the last one, given there are 3 factors, I implemented libfm in Keras and used the embeddings it computes.  I used between 3 and 5 latent factors.  All in all, these embeddings gave me a boost over 0.0010.  I think this is what led me in top 10.  I got some variety of models by varying which embeddings I was using.</li>\n</ul>\n\n<p><strong>Hyper parameter tuning</strong></p>\n\n<p>I spent time given how long it is to run an experiment, but I didn't tune much.  Main settings were to scale positive by around 400, use an initial score that minimizes expected loss if target is constant, and min child per leaf to be such that it requires at least 1 positive plus another example, in order to avoid overfiting to single positive examples.  I used 31 leaves and a depth of 8.  </p>\n\n<p><strong>Ensembling</strong></p>\n\n<p>On my local validation, the best way to blend several models was to average the logit of the predictions (aka raw predictions).  I started doing restacking, i.e. adding validation predictions to day 9 features, and training on it, but this was hitting my 50 feature limit, and runs were very long.  I did not ran it the last day for that reason.  It may have given me a little additional boost, but I don't think it would have been enough to move up in the LB, because my models were not diverse enough anyway.  I also think that using only day 9 for second level was leaving too much on the table.  I thought of generating of prediction for the full dataset, but that was a daunting task.  I now see this is what <a href=\"/bestfitting\">@bestfitting</a> did, great move on his part.</p>\n\n<p><strong>Takeaway</strong></p>\n\n<p>I think it was crazy to do this solo, too much work for a single person.  I really admire the other fools that went same way.  Once solo, I am not sure I should have done things differently, except for spending time to alleviate the 50 features limit.</p>\n\n<p>Also, as often in my competitions, I make a lot of progress the last day, not sure why.  In this case I moved from 0.9824 to 0.9828 public, and 0.9832 to 0.9837 private.  The lesson is to never give up, and not let the public LB dictate your mood.</p>\n\n<p>I hope the above will be useful to some.  Thanks for reading it all ;)</p>\n\n<p>Edit: I <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">shared my libFM implementation</a>.</p>",
  "messages": [
    {
      "id": 325218,
      "postDate": "2018-05-08T08:30:47.227Z",
      "content": "<p>First of all, all those who managed to get decent submissions out of this huge dataset deserve kudos.  Even if you score is below Kirk's shared kernel.  What you did is way more valuable, and lessons learned here will help you later.  </p>\n\n<p>Second, sharing is great when done in good faith, and lots of people did share a lot here, too many of them to name them.  Eve people who started here, like @Samrat, shared a lot.  This is what makes this community so valuable.</p>\n\n<p>Third, thanks to <a href=\"/inversion\">@inversion</a>, Kaggle, and Talking data for organizing a very challenging, ans almost leak free competition.  I say almost because it is clear now that the test data was sorted by click time then target value.  Exploiting this was the final twist that helped some of us fare better.  But the impact is not that large, I estimate it to be about 0.0004 for me.  And it was <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\">disclosed soon enough</a> for everyone to react to it.  Thanks to <a href=\"/plantsgo\">@plantsgo</a> for sharing it soon enough.</p>\n\n<p>I didn't decide to go solo from the start, but as time went by, I saw I was making progress every day, and decided to go solo till the end.  In retrospect I am not sure it was wise, I didn't sleep much in the last week ;)  I did receive some invites to merge during the last week before deadline, and I thank people for them.  I did miss a very late invite from a top 10 team as I was away that evening.  I wonder what would have been our score if we had teamed.</p>\n\n<p>Anyway here is my solution.  Given I was solo, and given the size of the data set, which meant hours to produce a submission, I decided to focus. I focused on a single type of model, LightGBM.  I split my time roughly as follows:</p>\n\n<ul>\n<li>80% feature engineering</li>\n<li>10% making local validation as fast as possible</li>\n<li>5% hyper parameter tuning</li>\n<li>5% ensembling</li>\n</ul>\n\n<p>I spent most of my time doing feature selection, as my machine was not usable with 50 features or more.  I wish I had used the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\">trick shared by Kruegger</a>, maybe I would have been able to add more features.  However, being forced to be selective about features probably led to better models in the end.  And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by <a href=\"/authman\">@authman</a>.</p>\n\n<p>My private LB score comes from a single lgb run with 48 features that scored 0.9825 public and 0.9835 private.  I submitted a blend of this with 5 other similar models that yield 0.9828 public and 0.9837 private, but for some reason <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56234\">that sub isn't taken into account</a>.  This is OK as my rank would not change with it.  I didn't submit all the 5 models individually during the competition, but did it after to get their score.  The best single lgb run scores 0.9827 public and 0.9836 private, with 48 features.</p>\n\n<p>I mostly used a 20 core Xeon at 2.3 MHz with 64 GB RAM and 64GB swap.  That machine is a bit slow, but it scales when using 20 threads.  I also used another machine with a 4 core i7 and 2 GPU to run Keras (see below).</p>\n\n<p><strong>Validation</strong></p>\n\n<p>Let's us look at validation.  This is key.  If you don't have a good validation scheme then you rely solely on LB probing, which can easily lead to overfit.  I ended up settling on:</p>\n\n<ul>\n<li>training on day &lt;=  8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.</li>\n<li>retraining on all data using 1.2 times the number of trees found by early stopping in validation</li>\n</ul>\n\n<p>Using two validation sets was to make sure I was not overfiting to one of them.  The hours were selected to match the public and private test data.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot. I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.</p>\n\n<p>I also used LB, i.e. only kept a feature if local auc and LB improved. Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.</p>\n\n<p>This was a very effective scheme, with the hour 4 score being the same as LB score with a std difference around 0.0001.  However, it is very time consuming because of the computation of the auc metric for early stopping.  In order to speed it for feature evaluation I used two lighter ways.  First, using only day 9 data, with 5% of hour 4 data for validation, the rest for training.  This could run in less than one hour, and was used as a filter.  Only features that improved on that went to the next stage which was train on day 8 and validate on day 9, both hour 4 and other test hours.  This was also a very effective scheme, with very good correlation with LB score, but it was not effective when evaluating lag features.  I therefore switched to training on day &lt;= 8 later on.</p>\n\n<p>Another way of speeding feature evaluation was to share each feature in a separate feather file.  This way, testing a feature set only requires assembling a set of files into one dataset.  Features were mostly tested by adding them one by one, and keeping them if local validation score improved by at least 0.00005.  I also added several of them at once, then removed them one by one to see if validation score decreased.  I basically did feature selection full time for the competition, preparing experiments to be run while I was away during day, or while I was sleeping.  The machine never stopped.</p>\n\n<p><strong>Feature Engineering</strong></p>\n\n<p>Features were computed on the concatenation of train and test_supplement, sorted by click time then by original order.  Now I am not sure the second item was useful.</p>\n\n<p>I used several families of features.</p>\n\n<ul>\n<li>Only  app, and os from the original features were kept. They were handled as categorical, and were my strongest 2 features with a third category made of the hour in the day.</li>\n<li>China days.  Introduced 24 periods that start at 4 pm.  These were\nused for lag features based on previous day(s) data.</li>\n<li>User: ip, device, os triplets.  </li>\n<li>Aggregates on various feature groups, similar to what was shared in many public kernels.  Aggregates I used were count, count of unique values, delta with previous value, delta with next value.  Time to next click when grouped by user was important.  Other useful ones I didn't see in kernels: delta with previous app.</li>\n<li>Lag features, based on previous China days values.  Previous count by some grouping, and previous target mean by some grouping.  The latter was a weighted average with the overall target mean, the weights being such that groups with few rows in it had a value closer to the overall average.  This is a standard normalization in target encoding.</li>\n<li>Ratios like number of clicks per ip, app to number of click per app.</li>\n<li>Not last.  This was to capture the leak.  It is one except for rows that are not the last of their group when grouped by user, app, and click time.  I ignored channel as I think that clicks are attributed to the most recent click having same user and app as the download.</li>\n<li>Target.  This is to also capture the leak.  I modified the target in train data by sorting is_attributed within group by user, app, and click time. The combination of both ways to capture the leak led to a boost between 0.0004 and 0.0005.</li>\n<li>Matrix factorization.  This was to capture the similarity between users and app.  I use several of them.  They all start with the same approach; construct a matrix with log of click counts. I used: ip x app, user x app, and os x device x app.  These matrices are extremely sparse (most values are 0).  For the first two I used truncated svd from sklearn, which gives me latent vectors (embeddings) for ip and user.  For the last one, given there are 3 factors, I implemented libfm in Keras and used the embeddings it computes.  I used between 3 and 5 latent factors.  All in all, these embeddings gave me a boost over 0.0010.  I think this is what led me in top 10.  I got some variety of models by varying which embeddings I was using.</li>\n</ul>\n\n<p><strong>Hyper parameter tuning</strong></p>\n\n<p>I spent time given how long it is to run an experiment, but I didn't tune much.  Main settings were to scale positive by around 400, use an initial score that minimizes expected loss if target is constant, and min child per leaf to be such that it requires at least 1 positive plus another example, in order to avoid overfiting to single positive examples.  I used 31 leaves and a depth of 8.  </p>\n\n<p><strong>Ensembling</strong></p>\n\n<p>On my local validation, the best way to blend several models was to average the logit of the predictions (aka raw predictions).  I started doing restacking, i.e. adding validation predictions to day 9 features, and training on it, but this was hitting my 50 feature limit, and runs were very long.  I did not ran it the last day for that reason.  It may have given me a little additional boost, but I don't think it would have been enough to move up in the LB, because my models were not diverse enough anyway.  I also think that using only day 9 for second level was leaving too much on the table.  I thought of generating of prediction for the full dataset, but that was a daunting task.  I now see this is what <a href=\"/bestfitting\">@bestfitting</a> did, great move on his part.</p>\n\n<p><strong>Takeaway</strong></p>\n\n<p>I think it was crazy to do this solo, too much work for a single person.  I really admire the other fools that went same way.  Once solo, I am not sure I should have done things differently, except for spending time to alleviate the 50 features limit.</p>\n\n<p>Also, as often in my competitions, I make a lot of progress the last day, not sure why.  In this case I moved from 0.9824 to 0.9828 public, and 0.9832 to 0.9837 private.  The lesson is to never give up, and not let the public LB dictate your mood.</p>\n\n<p>I hope the above will be useful to some.  Thanks for reading it all ;)</p>\n\n<p>Edit: I <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">shared my libFM implementation</a>.</p>",
      "rawMarkdown": "First of all, all those who managed to get decent submissions out of this huge dataset deserve kudos.  Even if you score is below Kirk's shared kernel.  What you did is way more valuable, and lessons learned here will help you later.  \n\nSecond, sharing is great when done in good faith, and lots of people did share a lot here, too many of them to name them.  Eve people who started here, like @Samrat, shared a lot.  This is what makes this community so valuable.\n\nThird, thanks to @inversion, Kaggle, and Talking data for organizing a very challenging, ans almost leak free competition.  I say almost because it is clear now that the test data was sorted by click time then target value.  Exploiting this was the final twist that helped some of us fare better.  But the impact is not that large, I estimate it to be about 0.0004 for me.  And it was [disclosed soon enough][1] for everyone to react to it.  Thanks to @plantsgo for sharing it soon enough.\n\nI didn't decide to go solo from the start, but as time went by, I saw I was making progress every day, and decided to go solo till the end.  In retrospect I am not sure it was wise, I didn't sleep much in the last week ;)  I did receive some invites to merge during the last week before deadline, and I thank people for them.  I did miss a very late invite from a top 10 team as I was away that evening.  I wonder what would have been our score if we had teamed.\n\nAnyway here is my solution.  Given I was solo, and given the size of the data set, which meant hours to produce a submission, I decided to focus. I focused on a single type of model, LightGBM.  I split my time roughly as follows:\n\n - 80% feature engineering\n - 10% making local validation as fast as possible\n - 5% hyper parameter tuning\n - 5% ensembling\n\nI spent most of my time doing feature selection, as my machine was not usable with 50 features or more.  I wish I had used the [trick shared by Kruegger][2], maybe I would have been able to add more features.  However, being forced to be selective about features probably led to better models in the end.  And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by @authman.\n\nMy private LB score comes from a single lgb run with 48 features that scored 0.9825 public and 0.9835 private.  I submitted a blend of this with 5 other similar models that yield 0.9828 public and 0.9837 private, but for some reason [that sub isn't taken into account][3].  This is OK as my rank would not change with it.  I didn't submit all the 5 models individually during the competition, but did it after to get their score.  The best single lgb run scores 0.9827 public and 0.9836 private, with 48 features.\n\nI mostly used a 20 core Xeon at 2.3 MHz with 64 GB RAM and 64GB swap.  That machine is a bit slow, but it scales when using 20 threads.  I also used another machine with a 4 core i7 and 2 GPU to run Keras (see below).\n\n**Validation**\n\nLet's us look at validation.  This is key.  If you don't have a good validation scheme then you rely solely on LB probing, which can easily lead to overfit.  I ended up settling on:\n\n - training on day &lt;=  8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.\n - retraining on all data using 1.2 times the number of trees found by early stopping in validation\n\nUsing two validation sets was to make sure I was not overfiting to one of them.  The hours were selected to match the public and private test data.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot. I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.\n\nI also used LB, i.e. only kept a feature if local auc and LB improved. Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.\n\n This was a very effective scheme, with the hour 4 score being the same as LB score with a std difference around 0.0001.  However, it is very time consuming because of the computation of the auc metric for early stopping.  In order to speed it for feature evaluation I used two lighter ways.  First, using only day 9 data, with 5% of hour 4 data for validation, the rest for training.  This could run in less than one hour, and was used as a filter.  Only features that improved on that went to the next stage which was train on day 8 and validate on day 9, both hour 4 and other test hours.  This was also a very effective scheme, with very good correlation with LB score, but it was not effective when evaluating lag features.  I therefore switched to training on day &lt;= 8 later on.\n\nAnother way of speeding feature evaluation was to share each feature in a separate feather file.  This way, testing a feature set only requires assembling a set of files into one dataset.  Features were mostly tested by adding them one by one, and keeping them if local validation score improved by at least 0.00005.  I also added several of them at once, then removed them one by one to see if validation score decreased.  I basically did feature selection full time for the competition, preparing experiments to be run while I was away during day, or while I was sleeping.  The machine never stopped.\n\n**Feature Engineering**\n\nFeatures were computed on the concatenation of train and test_supplement, sorted by click time then by original order.  Now I am not sure the second item was useful.\n\nI used several families of features.\n\n -  Only  app, and os from the original features were kept. They were handled as categorical, and were my strongest 2 features with a third category made of the hour in the day.\n - China days.  Introduced 24 periods that start at 4 pm.  These were\n   used for lag features based on previous day(s) data.\n - User: ip, device, os triplets.  \n - Aggregates on various feature groups, similar to what was shared in many public kernels.  Aggregates I used were count, count of unique values, delta with previous value, delta with next value.  Time to next click when grouped by user was important.  Other useful ones I didn't see in kernels: delta with previous app.\n - Lag features, based on previous China days values.  Previous count by some grouping, and previous target mean by some grouping.  The latter was a weighted average with the overall target mean, the weights being such that groups with few rows in it had a value closer to the overall average.  This is a standard normalization in target encoding.\n - Ratios like number of clicks per ip, app to number of click per app.\n - Not last.  This was to capture the leak.  It is one except for rows that are not the last of their group when grouped by user, app, and click time.  I ignored channel as I think that clicks are attributed to the most recent click having same user and app as the download.\n - Target.  This is to also capture the leak.  I modified the target in train data by sorting is_attributed within group by user, app, and click time. The combination of both ways to capture the leak led to a boost between 0.0004 and 0.0005.\n - Matrix factorization.  This was to capture the similarity between users and app.  I use several of them.  They all start with the same approach; construct a matrix with log of click counts. I used: ip x app, user x app, and os x device x app.  These matrices are extremely sparse (most values are 0).  For the first two I used truncated svd from sklearn, which gives me latent vectors (embeddings) for ip and user.  For the last one, given there are 3 factors, I implemented libfm in Keras and used the embeddings it computes.  I used between 3 and 5 latent factors.  All in all, these embeddings gave me a boost over 0.0010.  I think this is what led me in top 10.  I got some variety of models by varying which embeddings I was using.\n\n**Hyper parameter tuning**\n\nI spent time given how long it is to run an experiment, but I didn't tune much.  Main settings were to scale positive by around 400, use an initial score that minimizes expected loss if target is constant, and min child per leaf to be such that it requires at least 1 positive plus another example, in order to avoid overfiting to single positive examples.  I used 31 leaves and a depth of 8.  \n\n**Ensembling**\n\nOn my local validation, the best way to blend several models was to average the logit of the predictions (aka raw predictions).  I started doing restacking, i.e. adding validation predictions to day 9 features, and training on it, but this was hitting my 50 feature limit, and runs were very long.  I did not ran it the last day for that reason.  It may have given me a little additional boost, but I don't think it would have been enough to move up in the LB, because my models were not diverse enough anyway.  I also think that using only day 9 for second level was leaving too much on the table.  I thought of generating of prediction for the full dataset, but that was a daunting task.  I now see this is what @bestfitting did, great move on his part.\n\n**Takeaway**\n\nI think it was crazy to do this solo, too much work for a single person.  I really admire the other fools that went same way.  Once solo, I am not sure I should have done things differently, except for spending time to alleviate the 50 features limit.\n\nAlso, as often in my competitions, I make a lot of progress the last day, not sure why.  In this case I moved from 0.9824 to 0.9828 public, and 0.9832 to 0.9837 private.  The lesson is to never give up, and not let the public LB dictate your mood.\n\nI hope the above will be useful to some.  Thanks for reading it all ;)\n\nEdit: I [shared my libFM implementation][4].\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\n  [3]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56234\n  [4]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497",
      "votes": 238
    },
    {
      "id": 325525,
      "postDate": "2018-05-08T13:54:14.053Z",
      "content": "<p>Wow that's impressive. I love the matrix factorization part. Thanks for sharing in so much details (I wish I had access to your 250 messy files ....) Well done CPMP !</p>",
      "rawMarkdown": "Wow that's impressive. I love the matrix factorization part. Thanks for sharing in so much details (I wish I had access to your 250 messy files ....) Well done CPMP !",
      "votes": 7,
      "replies": [
        {
          "id": 325548,
          "postDate": "2018-05-08T14:33:32.937Z",
          "content": "<p>Thanks Olivier, well done too on your side too.  I sincerely hope that the pair you form with Yifan will get the gold it deserves in a coming competition.</p>",
          "rawMarkdown": "Thanks Olivier, well done too on your side too.  I sincerely hope that the pair you form with Yifan will get the gold it deserves in a coming competition.",
          "votes": 2
        },
        {
          "id": 325572,
          "postDate": "2018-05-08T15:06:32.453Z",
          "content": "<p>Thanks for your kind words. My team mates were amazing and Yifan did an oustanding job. I would be hundreds of places behind without them! Looking at the data size on coming competitions I'm afraid I won't be able to join the party ;-)</p>",
          "rawMarkdown": "Thanks for your kind words. My team mates were amazing and Yifan did an oustanding job. I would be hundreds of places behind without them! Looking at the data size on coming competitions I'm afraid I won't be able to join the party ;-)",
          "votes": 2
        }
      ]
    },
    {
      "id": 326842,
      "postDate": "2018-05-10T12:20:35.643Z",
      "content": "<p>I <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">shared my libFM implementation in Keras</a>.</p>",
      "rawMarkdown": "I [shared my libFM implementation in Keras][1].\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497",
      "votes": 4
    },
    {
      "id": 327712,
      "postDate": "2018-05-12T08:45:23.070Z",
      "content": "<p>Helpful indeed</p>",
      "rawMarkdown": "Helpful indeed",
      "votes": 2
    },
    {
      "id": 327834,
      "postDate": "2018-05-12T15:41:56.977Z",
      "content": "<p>Great writeup and well done for the 6th place finish!</p>",
      "rawMarkdown": "Great writeup and well done for the 6th place finish!",
      "votes": 1
    },
    {
      "id": 327721,
      "postDate": "2018-05-12T09:15:36.173Z",
      "content": "<p>very nice!!</p>",
      "rawMarkdown": "very nice!!",
      "votes": 1
    },
    {
      "id": 326647,
      "postDate": "2018-05-10T05:14:33.733Z",
      "content": "<p>@CPMP\nInteresting that you didn't use channel as categorical, in my first LGB model it was in top 3 predictors.</p>",
      "rawMarkdown": "@CPMP\nInteresting that you didn't use channel as categorical, in my first LGB model it was in top 3 predictors.",
      "votes": 1,
      "replies": [
        {
          "id": 326682,
          "postDate": "2018-05-10T07:14:30.163Z",
          "content": "<p>Hi, it was my top predictor according to lgb, but removing it did improve score.</p>",
          "rawMarkdown": "Hi, it was my top predictor according to lgb, but removing it did improve score.",
          "votes": 1
        },
        {
          "id": 326750,
          "postDate": "2018-05-10T09:14:34.917Z",
          "content": "<p>@CPMP\nThat's really surprising,\nI thought of feature importance as a guide to which features I keep, I never imagined taking out a top predictor and see if it improves accuracy!</p>\n\n<p>How did you come up with the idea?\nHow did you test the first features' importance?\nDid you train on all fields and took away 1 feature at a time and see which helped accuracy or not?\n(run all without ip, run all without channel. etc...)</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "@CPMP\nThat's really surprising,\nI thought of feature importance as a guide to which features I keep, I never imagined taking out a top predictor and see if it improves accuracy!\n\nHow did you come up with the idea?\nHow did you test the first features' importance?\nDid you train on all fields and took away 1 feature at a time and see which helped accuracy or not?\n(run all without ip, run all without channel. etc...)\n\n\nThanks!"
        },
        {
          "id": 326754,
          "postDate": "2018-05-10T09:20:24.763Z",
          "content": "<p>I ran quite a lot of experiments, maybe not 7,000 as Mamas did, but hundreds, probably around 500.  And removing original features was part of it.</p>",
          "rawMarkdown": "I ran quite a lot of experiments, maybe not 7,000 as Mamas did, but hundreds, probably around 500.  And removing original features was part of it."
        }
      ]
    },
    {
      "id": 326623,
      "postDate": "2018-05-10T04:04:28.647Z",
      "content": "<p>Congrats and thanks to your sharing. I think I study a lot from your solution.</p>",
      "rawMarkdown": "Congrats and thanks to your sharing. I think I study a lot from your solution.",
      "votes": 1,
      "replies": [
        {
          "id": 326755,
          "postDate": "2018-05-10T09:20:43.860Z",
          "content": "<p>Thanks, glad you find it interesting.</p>",
          "rawMarkdown": "Thanks, glad you find it interesting."
        }
      ]
    },
    {
      "id": 326423,
      "postDate": "2018-05-09T17:26:29.283Z",
      "content": "<p>@CPMP thanks for sharing. I feel like a pleb next to your effort.\nI've noticed from a few mistakes I've done even from your few notes.</p>",
      "rawMarkdown": "@CPMP thanks for sharing. I feel like a pleb next to your effort.\nI've noticed from a few mistakes I've done even from your few notes.\n",
      "votes": 1,
      "replies": [
        {
          "id": 326434,
          "postDate": "2018-05-09T17:39:49.920Z",
          "content": "<p>Thanks!  You did well too, I'd like to know how you get so much from ensembling given I basically failed at it.</p>",
          "rawMarkdown": "Thanks!  You did well too, I'd like to know how you get so much from ensembling given I basically failed at it.",
          "votes": 1
        },
        {
          "id": 326506,
          "postDate": "2018-05-09T20:57:05.667Z",
          "content": "<p>I figured out why I could improve so much.\nBecause I didn't have enough RAM (52GB including swap), I trained the data only on 2 days (8,9 &amp; 7,9) and in different run using  14 of the hours with 26 features. Because I chose each time a different subset of the dataset the correlation wasn't perfect and I could improve so much.</p>\n\n<p>I used weighted mean with automatic anomaly removal. weight was related to LB score and correlation between the different submission. higher score -&gt; higher weight, lower correlation -&gt; higher weight but a bit less drastic.</p>",
          "rawMarkdown": "I figured out why I could improve so much.\nBecause I didn't have enough RAM (52GB including swap), I trained the data only on 2 days (8,9 &amp; 7,9) and in different run using  14 of the hours with 26 features. Because I chose each time a different subset of the dataset the correlation wasn't perfect and I could improve so much.\n\nI used weighted mean with automatic anomaly removal. weight was related to LB score and correlation between the different submission. higher score -&gt; higher weight, lower correlation -&gt; higher weight but a bit less drastic.",
          "votes": 1
        },
        {
          "id": 326513,
          "postDate": "2018-05-09T21:05:39.833Z",
          "content": "<p>Thanks, makes sense now.</p>",
          "rawMarkdown": "Thanks, makes sense now."
        }
      ]
    },
    {
      "id": 326082,
      "postDate": "2018-05-09T08:41:26.667Z",
      "content": "<p>Hello，CPMP，thanks for your sharing。I am new to kaggle and I am not familiar with how to select features。When I construct a new feature ，how should I decide to add it or not？Just make a submission to see if it improves score？And when I construct some features，how should I decide they are good or not？Just make a submission to see if they improve score？Very thanks for your reply。</p>",
      "rawMarkdown": "Hello，CPMP，thanks for your sharing。I am new to kaggle and I am not familiar with how to select features。When I construct a new feature ，how should I decide to add it or not？Just make a submission to see if it improves score？And when I construct some features，how should I decide they are good or not？Just make a submission to see if they improve score？Very thanks for your reply。",
      "votes": 1,
      "replies": [
        {
          "id": 326086,
          "postDate": "2018-05-09T08:50:24.900Z",
          "content": "<blockquote>\n  <p>how should I decide to add it or not？</p>\n</blockquote>\n\n<p>Read again my writeup, I explain it.</p>\n\n<blockquote>\n  <p>Just make a submission to see if it improves score？</p>\n</blockquote>\n\n<p>Absolutely not.  That's what I called LB probing in y writeup.  This will make you feel good on public LB, and most probably result in a disaster on the private LB.  Key is to use local validation as I explain.</p>",
          "rawMarkdown": "&gt; how should I decide to add it or not？\n\nRead again my writeup, I explain it.\n\n&gt; Just make a submission to see if it improves score？\n\nAbsolutely not.  That's what I called LB probing in y writeup.  This will make you feel good on public LB, and most probably result in a disaster on the private LB.  Key is to use local validation as I explain."
        }
      ]
    },
    {
      "id": 326062,
      "postDate": "2018-05-09T08:08:54.390Z",
      "content": "<p>Congrats and Thanks! The MF part is really amazing!!!</p>",
      "rawMarkdown": "Congrats and Thanks! The MF part is really amazing!!!",
      "votes": 1,
      "replies": [
        {
          "id": 326112,
          "postDate": "2018-05-09T09:57:19.193Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 326023,
      "postDate": "2018-05-09T07:03:50.680Z",
      "content": "<p>By the way CPMP, how precious was intuitively the fact to use China day? </p>",
      "rawMarkdown": "By the way CPMP, how precious was intuitively the fact to use China day? ",
      "votes": 1,
      "replies": [
        {
          "id": 326055,
          "postDate": "2018-05-09T07:42:40.903Z",
          "content": "<p>lag features helped but not much, maybe 0.0002.</p>",
          "rawMarkdown": "lag features helped but not much, maybe 0.0002."
        }
      ]
    },
    {
      "id": 326010,
      "postDate": "2018-05-09T06:52:07.540Z",
      "content": "<p>Thanks for your sharing. Not finished reading your post yet!  But come here to comment with a big thumb up !  Very inspired about the \"Matrix factorization\" part !  </p>",
      "rawMarkdown": "Thanks for your sharing. Not finished reading your post yet!  But come here to comment with a big thumb up !  Very inspired about the \"Matrix factorization\" part !  ",
      "votes": 1
    },
    {
      "id": 325979,
      "postDate": "2018-05-09T06:32:13.207Z",
      "content": "<p>Hi CPMP, Do you have any tips/resources to get better at feature engineering? This is the hardest part for any data scientist and something that i am not currently very good at. </p>",
      "rawMarkdown": "Hi CPMP, Do you have any tips/resources to get better at feature engineering? This is the hardest part for any data scientist and something that i am not currently very good at. ",
      "votes": 1,
      "replies": [
        {
          "id": 325989,
          "postDate": "2018-05-09T06:40:24.477Z",
          "content": "<ol>\n<li>Have a reliable CV setting.  Without it you cannot do FE.</li>\n<li>Try a lot, fail a lot, learn from failures.</li>\n<li>Read top teams writeup after each competition, look for things you did not think of, learn from them.</li>\n</ol>",
          "rawMarkdown": "1.  Have a reliable CV setting.  Without it you cannot do FE.\n2. Try a lot, fail a lot, learn from failures.\n3. Read top teams writeup after each competition, look for things you did not think of, learn from them.",
          "votes": 1
        },
        {
          "id": 326016,
          "postDate": "2018-05-09T06:59:16.123Z",
          "content": "<p>Thanks for this, i will try to incorporate them religiously in the future. </p>\n\n<p>Are there any good resources as well to read up on FE? I usually read a few papers related to the topic and  try to embed some ideas from there. </p>",
          "rawMarkdown": "Thanks for this, i will try to incorporate them religiously in the future. \n\nAre there any good resources as well to read up on FE? I usually read a few papers related to the topic and  try to embed some ideas from there. "
        },
        {
          "id": 326056,
          "postDate": "2018-05-09T07:47:10.797Z",
          "content": "<p>Here are two great presentations:</p>\n\n<p><a href=\"https://www.slideshare.net/OwenZhang2/tips-for-data-science-competitions\">https://www.slideshare.net/OwenZhang2/tips-for-data-science-competitions</a></p>\n\n<p><a href=\"https://www.slideshare.net/HJvanVeen/feature-engineering-72376750\">https://www.slideshare.net/HJvanVeen/feature-engineering-72376750</a></p>",
          "rawMarkdown": "Here are two great presentations:\n\nhttps://www.slideshare.net/OwenZhang2/tips-for-data-science-competitions\n\nhttps://www.slideshare.net/HJvanVeen/feature-engineering-72376750",
          "votes": 2
        },
        {
          "id": 326222,
          "postDate": "2018-05-09T13:10:54.473Z",
          "content": "<p>Thanks and congrats by the way :P </p>",
          "rawMarkdown": "Thanks and congrats by the way :P ",
          "votes": 1
        }
      ]
    },
    {
      "id": 325975,
      "postDate": "2018-05-09T06:23:40.737Z",
      "content": "<p>Congrats CPMP and glad to see your writeup again.</p>\n\n<blockquote>\n  <p>China days. Introduced 24 periods that start at 4 pm. These were used\n  for lag features based on previous day(s) data.\n  Can you show some details and reasons about this feature ? </p>\n</blockquote>",
      "rawMarkdown": "Congrats CPMP and glad to see your writeup again.\n\n&gt; China days. Introduced 24 periods that start at 4 pm. These were used\n&gt; for lag features based on previous day(s) data.\nCan you show some details and reasons about this feature ? \n",
      "votes": 1,
      "replies": [
        {
          "id": 325978,
          "postDate": "2018-05-09T06:32:08.483Z",
          "content": "<p>They make test look like the rest of the data.  They have about the same number of clicks per day.  UTC days are imbalanced, first and last UTC days on the data have way less clicks, esp day 6.  </p>\n\n<p>Because day 6 is so small, lag features computed for day 7 are not very good.  </p>\n\n<p>Using China day means the first 24 hours of data have no lag, but then all remaining data has lag computed form at least a full day.  This yield better result in my model.</p>",
          "rawMarkdown": "They make test look like the rest of the data.  They have about the same number of clicks per day.  UTC days are imbalanced, first and last UTC days on the data have way less clicks, esp day 6.  \n\nBecause day 6 is so small, lag features computed for day 7 are not very good.  \n\nUsing China day means the first 24 hours of data have no lag, but then all remaining data has lag computed form at least a full day.  This yield better result in my model."
        },
        {
          "id": 326009,
          "postDate": "2018-05-09T06:51:46.727Z",
          "content": "<p>Thanks CPMP. It's a magic time handle.\nWill you share your code latter? Looking forward to seeing all your solutions especially matrix factorization.</p>",
          "rawMarkdown": "Thanks CPMP. It's a magic time handle.\nWill you share your code latter? Looking forward to seeing all your solutions especially matrix factorization.",
          "votes": -1
        },
        {
          "id": 327439,
          "postDate": "2018-05-11T14:58:41.103Z",
          "content": "<p>HI I shared my matrix factorization code.</p>",
          "rawMarkdown": "HI I shared my matrix factorization code."
        }
      ]
    },
    {
      "id": 325871,
      "postDate": "2018-05-09T01:48:17.173Z",
      "content": "<p>Hello, </p>\n\n<p>Thanks for sharing your very intuitive solution in a very detailed manner. As a novice in this type of competitions I really appreciate it and see it as an opportunity to learn from others. I have questions on feature engineering and forgive me if these are stupid questions :)</p>\n\n<ul>\n<li><p><strong>delta with previous/next value</strong>, e.g. <code>time to next click by user</code>, when you are calculating this feature how do you create it for test/val data. Aren't we missing time until next click information for test set. <code>time since previous click</code> is easier to understand but if we have a long time horizon in test set wouldn't it introduce noise since we might be over exaggerating time. These made sense when I looked at Rossmann competition where time since/until promotions were used but I am confused when target is involved.</p></li>\n<li><p>Similarly for lagged <strong>previous target mean</strong>, are we using a sliding window to calculate the target encodings, e.g. calculate target encoding on previous X days(regularized) to encode next Y days (not regularized).  Let's say for this case let val set be 5 hours of period, use 1 hour sliding window to estimate features for next 5 hours, is this correct interpretation ?</p></li>\n</ul>\n\n<p>Thanks Again !                                                                                                                                                 </p>",
      "rawMarkdown": "Hello, \n\nThanks for sharing your very intuitive solution in a very detailed manner. As a novice in this type of competitions I really appreciate it and see it as an opportunity to learn from others. I have questions on feature engineering and forgive me if these are stupid questions :)\n\n-  **delta with previous/next value**, e.g. `time to next click by user`, when you are calculating this feature how do you create it for test/val data. Aren't we missing time until next click information for test set. `time since previous click` is easier to understand but if we have a long time horizon in test set wouldn't it introduce noise since we might be over exaggerating time. These made sense when I looked at Rossmann competition where time since/until promotions were used but I am confused when target is involved.\n\n- Similarly for lagged **previous target mean**, are we using a sliding window to calculate the target encodings, e.g. calculate target encoding on previous X days(regularized) to encode next Y days (not regularized).  Let's say for this case let val set be 5 hours of period, use 1 hour sliding window to estimate features for next 5 hours, is this correct interpretation ?\n\nThanks Again !                                                                                                                                                 ",
      "votes": 1,
      "replies": [
        {
          "id": 325977,
          "postDate": "2018-05-09T06:29:39.783Z",
          "content": "<ol>\n<li><p>Features were computed on the concatenation of train and test_supplement.  test is a subset of test_supplement.  I did not compute features on test, I predict on test_supplement, then join.  I explained how I did it <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201\">here</a>.  Later, I added the not_last feature to the join columns.</p></li>\n<li><p>lagged features are computed on previous China day, not on a sliding window.</p></li>\n</ol>",
          "rawMarkdown": "1. Features were computed on the concatenation of train and test_supplement.  test is a subset of test_supplement.  I did not compute features on test, I predict on test_supplement, then join.  I explained how I did it [here][1].  Later, I added the not_last feature to the join columns.\n\n2. lagged features are computed on previous China day, not on a sliding window.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201",
          "votes": 2
        },
        {
          "id": 326048,
          "postDate": "2018-05-09T07:31:18.297Z",
          "content": "<p>Thanks !</p>",
          "rawMarkdown": "Thanks !"
        }
      ]
    },
    {
      "id": 325739,
      "postDate": "2018-05-08T19:52:51.537Z",
      "content": "<p>WOW,congrats and thanks for your writeup,it's really helpful.</p>",
      "rawMarkdown": "WOW,congrats and thanks for your writeup,it's really helpful.",
      "votes": 1,
      "replies": [
        {
          "id": 325742,
          "postDate": "2018-05-08T19:55:14.223Z",
          "content": "<p>Thanks, means a lot coming from you.  I wish i had thought of RNN as you used them!</p>",
          "rawMarkdown": "Thanks, means a lot coming from you.  I wish i had thought of RNN as you used them!",
          "votes": 2
        },
        {
          "id": 325749,
          "postDate": "2018-05-08T20:10:08.257Z",
          "content": "<p>A competition is a journey full of uncertainty,we can not try too much in limited time,we must make decision everyday,even every hour,so if we can prove something,find something interesting or learn something,we are successful.<br>I always treat a competition as a force to learn more,I read many papers and solutions about CTR at the beginning of this competition,so I am looking foward to your keras-FM.</p>",
          "rawMarkdown": "A competition is a journey full of uncertainty,we can not try too much in limited time,we must make decision everyday,even every hour,so if we can prove something,find something interesting or learn something,we are successful.<br>I always treat a competition as a force to learn more,I read many papers and solutions about CTR at the beginning of this competition,so I am looking foward to your keras-FM.",
          "votes": 5
        },
        {
          "id": 326148,
          "postDate": "2018-05-09T11:03:16.160Z",
          "content": "<p>I will publish the libfm part soon.</p>",
          "rawMarkdown": "I will publish the libfm part soon.",
          "votes": 2
        },
        {
          "id": 326844,
          "postDate": "2018-05-10T12:21:52.927Z",
          "content": "<blockquote>\n  <p>I will publish the libfm part soon.</p>\n</blockquote>\n\n<p>Done, see <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497</a></p>",
          "rawMarkdown": "&gt; I will publish the libfm part soon.\n\nDone, see https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497"
        }
      ]
    },
    {
      "id": 325692,
      "postDate": "2018-05-08T18:13:16.287Z",
      "content": "<p>CPMP, Congratulations! Thanks for sharing highlights of your solution. Your contributions to the Kaggle community are indeed remarkable and highly appreciated. The matrix factorization idea is ingenious. I agree with you that having a good CV methodology is key esp. for a competition with so much data. I also used three setups (small/medium/large) that successively validated my results. But I can see numerous other things I did not do and your post gives me so many ideas for the next set of competitions :-) So thanks again.</p>",
      "rawMarkdown": "CPMP, Congratulations! Thanks for sharing highlights of your solution. Your contributions to the Kaggle community are indeed remarkable and highly appreciated. The matrix factorization idea is ingenious. I agree with you that having a good CV methodology is key esp. for a competition with so much data. I also used three setups (small/medium/large) that successively validated my results. But I can see numerous other things I did not do and your post gives me so many ideas for the next set of competitions :-) So thanks again.",
      "votes": 1,
      "replies": [
        {
          "id": 325744,
          "postDate": "2018-05-08T19:55:47.830Z",
          "content": "<p>Thanks a lot Adarsh!</p>",
          "rawMarkdown": "Thanks a lot Adarsh!",
          "votes": 1
        }
      ]
    },
    {
      "id": 325643,
      "postDate": "2018-05-08T16:57:11.780Z",
      "content": "<p>Thanks a lot for this detailed explanation. I was disappointed like a lot of other people to see my score dropped 400 place in 1h, and see people with 2 submissions beat me while I did hundreds of tries. </p>\n\n<p>But this was not the reason why I joined this competition, and it's definitely the high quality explanations / advises that you and others Kagglers provided that made me enjoyed the learning process. So thank you for that ! </p>",
      "rawMarkdown": "Thanks a lot for this detailed explanation. I was disappointed like a lot of other people to see my score dropped 400 place in 1h, and see people with 2 submissions beat me while I did hundreds of tries. \n\nBut this was not the reason why I joined this competition, and it's definitely the high quality explanations / advises that you and others Kagglers provided that made me enjoyed the learning process. So thank you for that ! ",
      "votes": 1
    },
    {
      "id": 325600,
      "postDate": "2018-05-08T16:06:16.107Z",
      "content": "<p>Thanks for sharing, I need to implement your feature selection strategies next time and practice getting rid of the features that took hours to engineer :)</p>\n\n<p>By the way, the ratio of the number of clicks per ip, day to the number of clicks per day was also working really well for me. Always in the top 15 features based on the feature importance. </p>",
      "rawMarkdown": "Thanks for sharing, I need to implement your feature selection strategies next time and practice getting rid of the features that took hours to engineer :)\n\nBy the way, the ratio of the number of clicks per ip, day to the number of clicks per day was also working really well for me. Always in the top 15 features based on the feature importance. ",
      "votes": 1
    },
    {
      "id": 325487,
      "postDate": "2018-05-08T13:16:26.200Z",
      "content": "<p>Some question about the lag features:if the day is 7,the only data that you can use is data in the day 6,but the data amount in the day 6 is very small,this mean your lag feature in the day 7 is almost 0??</p>",
      "rawMarkdown": "Some question about the lag features:if the day is 7,the only data that you can use is data in the day 6,but the data amount in the day 6 is very small,this mean your lag feature in the day 7 is almost 0??",
      "votes": 1,
      "replies": [
        {
          "id": 325502,
          "postDate": "2018-05-08T13:35:28.583Z",
          "content": "<p>Great question.  Lag features are undefined for first China day.</p>",
          "rawMarkdown": "Great question.  Lag features are undefined for first China day.",
          "votes": 1
        }
      ]
    },
    {
      "id": 325456,
      "postDate": "2018-05-08T12:41:27.470Z",
      "content": "<p>great lesson learnt here is implementing matrix factorization for categorical combinations with large number of levels. I think I used most other methods similarly in this post, but didn't do the libfm, which makes the difference.  </p>\n\n<p>Definitely should learn that, and hopefully to use it in future competition and datasets.  Do you have any code base/implementation/method paper reference that we could look into? Thanks!!</p>\n\n<p>Again, thanks for great sharing and always something useful learnt here. </p>",
      "rawMarkdown": "great lesson learnt here is implementing matrix factorization for categorical combinations with large number of levels. I think I used most other methods similarly in this post, but didn't do the libfm, which makes the difference.  \n\nDefinitely should learn that, and hopefully to use it in future competition and datasets.  Do you have any code base/implementation/method paper reference that we could look into? Thanks!!\n\nAgain, thanks for great sharing and always something useful learnt here. ",
      "votes": 1,
      "replies": [
        {
          "id": 325477,
          "postDate": "2018-05-08T13:01:28.050Z",
          "content": "<p>Start with <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html\">http://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html</a> and <a href=\"http://scikit-learn.org/stable/modules/decomposition.html\">http://scikit-learn.org/stable/modules/decomposition.html</a></p>\n\n<p>Truncated SVD is similar to PCA.</p>\n\n<p>For libfm, I'll write something on my blog on how to implement it in Keras.</p>",
          "rawMarkdown": "Start with http://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html and http://scikit-learn.org/stable/modules/decomposition.html\n\nTruncated SVD is similar to PCA.\n\nFor libfm, I'll write something on my blog on how to implement it in Keras.",
          "votes": 3
        },
        {
          "id": 325489,
          "postDate": "2018-05-08T13:17:50.460Z",
          "content": "<p>What is the link to the blog? :-)</p>",
          "rawMarkdown": "What is the link to the blog? :-)"
        },
        {
          "id": 325501,
          "postDate": "2018-05-08T13:33:54.437Z",
          "content": "<p>@CPMP,  Thanks!! @Olecram, this is CPMP's blog,<a href=\"https://www.ibm.com/developerworks/community/blogs/jfp?lang=en\">https://www.ibm.com/developerworks/community/blogs/jfp?lang=en</a> \nanother good resource to learn. </p>",
          "rawMarkdown": "@CPMP,  Thanks!! @Olecram, this is CPMP's blog,https://www.ibm.com/developerworks/community/blogs/jfp?lang=en \nanother good resource to learn. ",
          "votes": 3
        },
        {
          "id": 325522,
          "postDate": "2018-05-08T13:50:59.437Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 325450,
      "postDate": "2018-05-08T12:31:23.623Z",
      "content": "<p>Congratulations and thanks @CPMP for sharing your solution. </p>\n\n<p>I have learnt a few things from your FE i.e. the matrix factorization and the features you created to capture the leak. I now know my best single model with 21 features is nowhere near enough if you are talking 50 features :-)</p>\n\n<p>You are right about it being crazy to do this one solo. I decided to do it solo because this is my 1st pure time series competition hence; 1) I wanted to focus mainly on the learning part, 2) I felt like I do not have a lot of time hence may not be able to carry my weight to satisfaction if I joined a team. But no regrets as I have learnt so much.</p>",
      "rawMarkdown": "Congratulations and thanks @CPMP for sharing your solution. \n\nI have learnt a few things from your FE i.e. the matrix factorization and the features you created to capture the leak. I now know my best single model with 21 features is nowhere near enough if you are talking 50 features :-)\n\nYou are right about it being crazy to do this one solo. I decided to do it solo because this is my 1st pure time series competition hence; 1) I wanted to focus mainly on the learning part, 2) I felt like I do not have a lot of time hence may not be able to carry my weight to satisfaction if I joined a team. But no regrets as I have learnt so much.",
      "votes": 1,
      "replies": [
        {
          "id": 325478,
          "postDate": "2018-05-08T13:02:13.803Z",
          "content": "<p>Bravo for you solo, sorry your rank was hurt by Dirk's late share.</p>",
          "rawMarkdown": "Bravo for you solo, sorry your rank was hurt by Dirk's late share.",
          "votes": 1
        },
        {
          "id": 325492,
          "postDate": "2018-05-08T13:21:33.500Z",
          "content": "<p>Thanks. Although I have only been doing this for roughly one year, I have learnt to expect such surprises in the last days of a competition. So I just shake it off and move on. The real surprise to me is that a blend of blend average ending up at #153 on the private LB with a silver medal.</p>",
          "rawMarkdown": "Thanks. Although I have only been doing this for roughly one year, I have learnt to expect such surprises in the last days of a competition. So I just shake it off and move on. The real surprise to me is that a blend of blend average ending up at #153 on the private LB with a silver medal."
        }
      ]
    },
    {
      "id": 325365,
      "postDate": "2018-05-08T10:53:52.403Z",
      "content": "<p>You really did a great job! Thx for sharing. You can have a good dream today. ;)</p>",
      "rawMarkdown": "You really did a great job! Thx for sharing. You can have a good dream today. ;)",
      "votes": 1
    },
    {
      "id": 325361,
      "postDate": "2018-05-08T10:47:50.360Z",
      "content": "<p>Hi CPMP, thanks a lot for insightful sharing. You and top kagglers keep us up to the game as you guys showed that smarter ways can win the game, not just copy/blend others results. </p>",
      "rawMarkdown": "Hi CPMP, thanks a lot for insightful sharing. You and top kagglers keep us up to the game as you guys showed that smarter ways can win the game, not just copy/blend others results. ",
      "votes": 1
    },
    {
      "id": 325313,
      "postDate": "2018-05-08T09:45:46.773Z",
      "content": "<p>Well done CPMP</p>",
      "rawMarkdown": "Well done CPMP",
      "votes": 1
    },
    {
      "id": 325310,
      "postDate": "2018-05-08T09:42:20.970Z",
      "content": "<p>@CPMP Nice work!. I will repeat your results when i am free :)</p>",
      "rawMarkdown": "@CPMP Nice work!. I will repeat your results when i am free :)",
      "votes": 1,
      "replies": [
        {
          "id": 325421,
          "postDate": "2018-05-08T11:57:40.117Z",
          "content": "<p>Thanks!  Let me know if and when you start it, would be interested to see how it goes.</p>",
          "rawMarkdown": "Thanks!  Let me know if and when you start it, would be interested to see how it goes."
        }
      ]
    },
    {
      "id": 325242,
      "postDate": "2018-05-08T08:55:44.717Z",
      "content": "<p>Hi CPMP, well done and thanks for this fruitful insights!</p>",
      "rawMarkdown": "Hi CPMP, well done and thanks for this fruitful insights!",
      "votes": 1,
      "replies": [
        {
          "id": 325263,
          "postDate": "2018-05-08T09:06:55.833Z",
          "content": "<p>I like your comment: my machine is a little bit slow! It is a <strong>20 core</strong> Xeon at 2.3 MHz with <strong>64 GB RAM</strong> and 64GB swap. This is way beyond my machine. And thanks for these nice details. Very interesting. Kaggle rocks and keeps me up at night!</p>",
          "rawMarkdown": "I like your comment: my machine is a little bit slow! It is a **20 core** Xeon at 2.3 MHz with **64 GB RAM** and 64GB swap. This is way beyond my machine. And thanks for these nice details. Very interesting. Kaggle rocks and keeps me up at night!",
          "votes": 2
        },
        {
          "id": 325269,
          "postDate": "2018-05-08T09:13:02.863Z",
          "content": "<p>Thanks!</p>\n\n<p>regarding our point on machines, any VM on a cloud is way faster.  A 4 cores i7 with SSD is faster.  A Macbook pro is faster.</p>\n\n<p>But yes, this is bigger and faster than an entry level laptop.  </p>\n\n<p>I used to compete on Kaggle with a macbook pro with 16GB.  I know how challenging it is to work with large datasets...  Net result is that I fried my macbookpro and decided to buy a relatively cheap machine that could scale out.  Yu can find second hand machines similar to it at a very reasonable price now.</p>",
          "rawMarkdown": "Thanks!\n\nregarding our point on machines, any VM on a cloud is way faster.  A 4 cores i7 with SSD is faster.  A Macbook pro is faster.\n\nBut yes, this is bigger and faster than an entry level laptop.  \n\nI used to compete on Kaggle with a macbook pro with 16GB.  I know how challenging it is to work with large datasets...  Net result is that I fried my macbookpro and decided to buy a relatively cheap machine that could scale out.  Yu can find second hand machines similar to it at a very reasonable price now."
        },
        {
          "id": 325289,
          "postDate": "2018-05-08T09:24:49.847Z",
          "content": "<p>For the 2nd half of the competition I was using a core i5 machine with 16GB RAM and 100+GB Swap space on 2 SSD's... I just shut down my machine after many weeks... Luckily I had an access to a better system in the first half and thats when I had quick and promising results..</p>",
          "rawMarkdown": "For the 2nd half of the competition I was using a core i5 machine with 16GB RAM and 100+GB Swap space on 2 SSD's... I just shut down my machine after many weeks... Luckily I had an access to a better system in the first half and thats when I had quick and promising results..",
          "votes": 1
        },
        {
          "id": 328627,
          "postDate": "2018-05-14T18:38:01.220Z",
          "content": "<p>What OS  on machines, you folks are discussing? Linux/Window/MAC? I see only afordable Window machines for 32/64G in the market.  Linux very few and expensive.</p>",
          "rawMarkdown": "What OS  on machines, you folks are discussing? Linux/Window/MAC? I see only afordable Window machines for 32/64G in the market.  Linux very few and expensive."
        },
        {
          "id": 328673,
          "postDate": "2018-05-14T21:29:56.677Z",
          "content": "<p>You can always install Linux on a Windows machine.  </p>",
          "rawMarkdown": "You can always install Linux on a Windows machine.  "
        },
        {
          "id": 337640,
          "postDate": "2018-06-03T11:54:14.127Z",
          "content": "<p>Yes , I have just installed kali</p>",
          "rawMarkdown": "Yes , I have just installed kali",
          "votes": 1
        }
      ]
    },
    {
      "id": 325238,
      "postDate": "2018-05-08T08:52:36.390Z",
      "content": "<p>Thanks a ton for sharing this info <a href=\"/cpmpml\">@cpmpml</a> BTW do you plan to share any code, may be via github.. And thanks a lot for the mention tooo :-)</p>",
      "rawMarkdown": "Thanks a ton for sharing this info @cpmpml BTW do you plan to share any code, may be via github.. And thanks a lot for the mention tooo :-)",
      "votes": 1,
      "replies": [
        {
          "id": 325241,
          "postDate": "2018-05-08T08:55:21.057Z",
          "content": "<p>I won't share the code in whole as it is messy, I didn't use source code control, rather copied files around. I have 250 files in my model directory...  Cleaning this and documenting is a daunting task, I am glad to not be in prize position for that ;)</p>\n\n<p>But I will share some of it via my blog.  I'll put pointer to it from here when I do it.</p>",
          "rawMarkdown": "I won't share the code in whole as it is messy, I didn't use source code control, rather copied files around. I have 250 files in my model directory...  Cleaning this and documenting is a daunting task, I am glad to not be in prize position for that ;)\n\nBut I will share some of it via my blog.  I'll put pointer to it from here when I do it.",
          "votes": 4
        },
        {
          "id": 325244,
          "postDate": "2018-05-08T08:56:16.047Z",
          "content": "<p>CPMP, where is exactly your blog?</p>",
          "rawMarkdown": "CPMP, where is exactly your blog?",
          "votes": 1
        },
        {
          "id": 325282,
          "postDate": "2018-05-08T09:19:27.250Z",
          "content": "<p>Here: <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp?lang=en\">https://www.ibm.com/developerworks/community/blogs/jfp?lang=en</a></p>\n\n<p>Not very active lately because of Kaggle.</p>",
          "rawMarkdown": "Here: https://www.ibm.com/developerworks/community/blogs/jfp?lang=en\n\nNot very active lately because of Kaggle.",
          "votes": 3
        },
        {
          "id": 325284,
          "postDate": "2018-05-08T09:21:05.710Z",
          "content": "<p>Yeah snippets of code should do....</p>",
          "rawMarkdown": "Yeah snippets of code should do....",
          "votes": 1
        },
        {
          "id": 326021,
          "postDate": "2018-05-09T07:02:43.613Z",
          "content": "<p>Thanks CPMP for the link about your blog. I will definitely keep an eye on it! And thank you in general for being so talkative on kaggle forum. You are my hero! No wonder why you are a discussion grand master and ranked n°1.. Keep it up. It is higly appreciated!</p>",
          "rawMarkdown": "Thanks CPMP for the link about your blog. I will definitely keep an eye on it! And thank you in general for being so talkative on kaggle forum. You are my hero! No wonder why you are a discussion grand master and ranked n°1.. Keep it up. It is higly appreciated!",
          "votes": 1
        }
      ]
    },
    {
      "id": 325235,
      "postDate": "2018-05-08T08:47:05.573Z",
      "content": "<p>Well done - a bucket full of awesome</p>",
      "rawMarkdown": "Well done - a bucket full of awesome",
      "votes": 1,
      "replies": [
        {
          "id": 325445,
          "postDate": "2018-05-08T12:25:58.697Z",
          "content": "<p>You did well too, congrats!</p>",
          "rawMarkdown": "You did well too, congrats!",
          "votes": 1
        }
      ]
    },
    {
      "id": 330163,
      "postDate": "2018-05-18T07:05:52.493Z",
      "content": "<p>Thank you very much for sharing the detailed information, <a href=\"/cpmpml\">@cpmpml</a>.</p>\n\n<blockquote>\n  <p>And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by @autheman.</p>\n</blockquote>\n\n<p>What are past40 features and 'two_round' parameter?</p>",
      "rawMarkdown": "Thank you very much for sharing the detailed information, @cpmpml.\n\n&gt; And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by @autheman.\n\nWhat are past40 features and 'two_round' parameter?\n",
      "votes": 2,
      "replies": [
        {
          "id": 330231,
          "postDate": "2018-05-18T10:26:42.147Z",
          "content": "<p>I would not have been able to use more than 40 parameters if I had not set the 'two_round' lightgbm parameter to True.</p>",
          "rawMarkdown": "I would not have been able to use more than 40 parameters if I had not set the 'two_round' lightgbm parameter to True.",
          "votes": 2
        }
      ]
    },
    {
      "id": 329923,
      "postDate": "2018-05-17T15:05:31.850Z",
      "content": "<p>Thanks for sharing! Very interesting for matrix factorization part!</p>",
      "rawMarkdown": "Thanks for sharing! Very interesting for matrix factorization part!",
      "votes": 2
    },
    {
      "id": 329192,
      "postDate": "2018-05-16T00:34:56.440Z",
      "content": "<p>That very helpful</p>",
      "rawMarkdown": "That very helpful",
      "votes": 2
    },
    {
      "id": 328884,
      "postDate": "2018-05-15T09:08:33.700Z",
      "content": "<p>very helpful!</p>",
      "rawMarkdown": "very helpful!",
      "votes": 2
    },
    {
      "id": 328407,
      "postDate": "2018-05-14T08:28:52.663Z",
      "content": "<p>really helpful!</p>",
      "rawMarkdown": "really helpful!",
      "votes": 2
    },
    {
      "id": 327610,
      "postDate": "2018-05-12T01:58:57.153Z",
      "content": "<p>so cool</p>",
      "rawMarkdown": "so cool",
      "votes": 2
    },
    {
      "id": 327606,
      "postDate": "2018-05-12T01:34:17.413Z",
      "content": "<p>Great work</p>",
      "rawMarkdown": "Great work",
      "votes": 2
    },
    {
      "id": 327399,
      "postDate": "2018-05-11T12:55:59.870Z",
      "content": "<p>good</p>",
      "rawMarkdown": "good",
      "votes": 2
    },
    {
      "id": 327391,
      "postDate": "2018-05-11T12:38:10.763Z",
      "content": "<p>Great work@CPMP! Congratulations and thanks for detailed sharing, your tips are really useful!</p>",
      "rawMarkdown": "Great work@CPMP! Congratulations and thanks for detailed sharing, your tips are really useful!",
      "votes": 2
    },
    {
      "id": 327235,
      "postDate": "2018-05-11T04:51:09.847Z",
      "content": "<p>Hello all. I'm brand new around here, having just completed the 'my first model' training in the learn section. Look forward to the day when I can leave substantive comments because I understand what y'all are talking about. :)</p>",
      "rawMarkdown": "Hello all. I'm brand new around here, having just completed the 'my first model' training in the learn section. Look forward to the day when I can leave substantive comments because I understand what y'all are talking about. :)",
      "votes": 2
    },
    {
      "id": 326672,
      "postDate": "2018-05-10T06:43:08.703Z",
      "content": "<p>Thank your for your sharing @CPMP. I also did matrix factorization, using ip x app (my machine can't handle the size of user x app) with count of clicks instead of log counts you used. My procedure is to use ALS to extract 4 or 8 factors corresponding to ip, however both set of factors doesn't improve my validation score and I'm wondering whether it is a problem of my way of counting, ALS, or using only ip x app.  Please share your thoughts with me.</p>",
      "rawMarkdown": "Thank your for your sharing @CPMP. I also did matrix factorization, using ip x app (my machine can't handle the size of user x app) with count of clicks instead of log counts you used. My procedure is to use ALS to extract 4 or 8 factors corresponding to ip, however both set of factors doesn't improve my validation score and I'm wondering whether it is a problem of my way of counting, ALS, or using only ip x app.  Please share your thoughts with me.",
      "votes": 2,
      "replies": [
        {
          "id": 326733,
          "postDate": "2018-05-10T08:42:07.393Z",
          "content": "<p>Hi, I found that adding the app factors in this case was hurting, I only kept the ip factors (and the user factors with the other matrix factorization).  I don't know why using app factors is hurting.</p>\n\n<p>Using ALS is fine I think.</p>",
          "rawMarkdown": "Hi, I found that adding the app factors in this case was hurting, I only kept the ip factors (and the user factors with the other matrix factorization).  I don't know why using app factors is hurting.\n\nUsing ALS is fine I think."
        },
        {
          "id": 326751,
          "postDate": "2018-05-10T09:16:55.490Z",
          "content": "<p>@CPMP\nCould you share a source where i can read about the matrix factorization thingy that you used?\nI don't want to use much of your time explaining, i tried to find something online but it obviously contains information only on the calculation method itself.\nI didn't understand the actual numeric outcome of the matrix factorization you did. what its purpose is...</p>",
          "rawMarkdown": "@CPMP\nCould you share a source where i can read about the matrix factorization thingy that you used?\nI don't want to use much of your time explaining, i tried to find something online but it obviously contains information only on the calculation method itself.\nI didn't understand the actual numeric outcome of the matrix factorization you did. what its purpose is..."
        },
        {
          "id": 326752,
          "postDate": "2018-05-10T09:18:26.207Z",
          "content": "<p>I had the idea myself, no external source. I'm going to publish soon my libfm in keras and how it is used here.</p>",
          "rawMarkdown": "I had the idea myself, no external source. I'm going to publish soon my libfm in keras and how it is used here.",
          "votes": 1
        },
        {
          "id": 326818,
          "postDate": "2018-05-10T11:38:02.257Z",
          "content": "<p>Well I didn't mention it clearly, I kept and used only ip factors and the score doesn't improved. I'm guessing that it is a result of using raw counts instead of log counts, or maybe user factors have more predictive power than ip factors. Thank you very much anyway, I learned a lot from your solution!</p>",
          "rawMarkdown": "Well I didn't mention it clearly, I kept and used only ip factors and the score doesn't improved. I'm guessing that it is a result of using raw counts instead of log counts, or maybe user factors have more predictive power than ip factors. Thank you very much anyway, I learned a lot from your solution!"
        },
        {
          "id": 326826,
          "postDate": "2018-05-10T11:44:29.893Z",
          "content": "<p>@AmirH you can get a flavor of MF from <a href=\"https://beckernick.github.io/matrix-factorization-recommender/\">this blogpost</a> and <a href=\"https://www.coursera.org/learn/competitive-data-science/lecture/8o1Hc/matrix-factorizations\">this video lecture</a>.<br> But I will suggest you to wait for CPMP blog post.</p>",
          "rawMarkdown": "@AmirH you can get a flavor of MF from [this blogpost][1] and [this video lecture][2].<br> But I will suggest you to wait for CPMP blog post.\n\n\n  [1]: https://beckernick.github.io/matrix-factorization-recommender/\n  [2]: https://www.coursera.org/learn/competitive-data-science/lecture/8o1Hc/matrix-factorizations",
          "votes": 2
        },
        {
          "id": 327028,
          "postDate": "2018-05-10T17:42:23.867Z",
          "content": "<p>Thanks @CPMP and @Sohaib Omar!!!</p>",
          "rawMarkdown": "Thanks @CPMP and @Sohaib Omar!!!",
          "votes": 1
        },
        {
          "id": 327031,
          "postDate": "2018-05-10T17:53:59.397Z",
          "content": "<p>@Jimmy Chen, raw counts are not good at all.</p>",
          "rawMarkdown": "@Jimmy Chen, raw counts are not good at all."
        }
      ]
    },
    {
      "id": 326089,
      "postDate": "2018-05-09T09:01:01.820Z",
      "content": "<p>Thanks a lot fortaking the time to explain and share your solutions and insight ! I'm not sure I fully understand the matrix factorization part, which seems really interesting. Do you have any detailed explanation of this ? (article, code, anything really ?)</p>",
      "rawMarkdown": "Thanks a lot fortaking the time to explain and share your solutions and insight ! I'm not sure I fully understand the matrix factorization part, which seems really interesting. Do you have any detailed explanation of this ? (article, code, anything really ?)",
      "votes": 2,
      "replies": [
        {
          "id": 326093,
          "postDate": "2018-05-09T09:04:20.823Z",
          "content": "<p>I'll publish the libfm in keras part, hopefully it will answer your question.</p>",
          "rawMarkdown": "I'll publish the libfm in keras part, hopefully it will answer your question.",
          "votes": 4
        },
        {
          "id": 326124,
          "postDate": "2018-05-09T10:17:14.083Z",
          "content": "<p>Thanks, let me know !</p>",
          "rawMarkdown": "Thanks, let me know !"
        },
        {
          "id": 326843,
          "postDate": "2018-05-10T12:21:18.190Z",
          "content": "<p>Done, see <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497</a></p>",
          "rawMarkdown": "Done, see https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497",
          "votes": 2
        },
        {
          "id": 327258,
          "postDate": "2018-05-11T06:34:26.533Z",
          "content": "<p>Thanks for tagging me !</p>",
          "rawMarkdown": "Thanks for tagging me !"
        }
      ]
    },
    {
      "id": 325484,
      "postDate": "2018-05-08T13:11:40.677Z",
      "content": "<p>Very detailed description, I read it carefully and it contains many golden experiences. I must say thank you for your endevor. </p>\n\n<p>BTW, there's one thing I would like to know a bit more: you mentioned that you used 1.2 times num_boost_round found by early stopping on validation set. How is 1.2 determined? Is that determined by the ratio of samples on training set for local validation over training set for submission?</p>",
      "rawMarkdown": "Very detailed description, I read it carefully and it contains many golden experiences. I must say thank you for your endevor. \n\nBTW, there's one thing I would like to know a bit more: you mentioned that you used 1.2 times num_boost_round found by early stopping on validation set. How is 1.2 determined? Is that determined by the ratio of samples on training set for local validation over training set for submission?",
      "votes": 2,
      "replies": [
        {
          "id": 325504,
          "postDate": "2018-05-08T13:36:24.807Z",
          "content": "<p>I tried 1.0, 1.1, 1.2, and 1.3 on two different models, and 1.2 was giving the best LB result on both.</p>",
          "rawMarkdown": "I tried 1.0, 1.1, 1.2, and 1.3 on two different models, and 1.2 was giving the best LB result on both.",
          "votes": 2
        },
        {
          "id": 325517,
          "postDate": "2018-05-08T13:47:15.443Z",
          "content": "<p>Thanks! It seems it really needs some trials.</p>",
          "rawMarkdown": "Thanks! It seems it really needs some trials."
        },
        {
          "id": 325520,
          "postDate": "2018-05-08T13:49:13.390Z",
          "content": "<p>Usually people use between 1% and 15% more, but here it seems that going further helps.  I would not draw a genera conclusion though.</p>",
          "rawMarkdown": "Usually people use between 1% and 15% more, but here it seems that going further helps.  I would not draw a genera conclusion though."
        }
      ]
    },
    {
      "id": 1316940,
      "postDate": "2021-05-21T03:31:38.053Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>I saw here you were using a slow computer for your work, would your results or work process be better if you had more GPU's ?</p>",
      "rawMarkdown": "Hi @cpmpml \n\nI saw here you were using a slow computer for your work, would your results or work process be better if you had more GPU's ?"
    },
    {
      "id": 408114,
      "postDate": "2018-10-22T10:38:36.190Z",
      "content": "<p>Thank you for sharing this Sir.It helped me a lot....</p>",
      "rawMarkdown": "Thank you for sharing this Sir.It helped me a lot...."
    },
    {
      "id": 333126,
      "postDate": "2018-05-24T12:45:13.913Z",
      "content": "<blockquote>\n  <p>and min child per leaf to be such that it requires at least 1 positive plus another example</p>\n</blockquote>\n\n<p>How did you achieve that <a href=\"/cpmpml\">@cpmpml</a> ? parameter value(which) or modifying source code? Whould you mind sharing?\nThanks in advance and congrats</p>",
      "rawMarkdown": " \n\n&gt; and min child per leaf to be such that it requires at least 1 positive plus another example\n\nHow did you achieve that @cpmpml ? parameter value(which) or modifying source code? Whould you mind sharing?\nThanks in advance and congrats",
      "replies": [
        {
          "id": 334239,
          "postDate": "2018-05-26T18:50:32.083Z",
          "content": "<p>I set <code>min_data_in_leaf</code> to 1 + <code>scale_pos_weight</code> value.</p>",
          "rawMarkdown": "I set `min_data_in_leaf` to 1 + `scale_pos_weight` value.",
          "votes": 1
        },
        {
          "id": 339663,
          "postDate": "2018-06-07T09:56:59.307Z",
          "content": "<p>Are you sure that min_data_in_leaf takes scale_pos_weight into account?\n\"min_child_weight \" parameter might be more appropriate? but I am not sure about the \"hessian\" value ...</p>",
          "rawMarkdown": "Are you sure that min_data_in_leaf takes scale_pos_weight into account?\n\"min_child_weight \" parameter might be more appropriate? but I am not sure about the \"hessian\" value ..."
        },
        {
          "id": 339666,
          "postDate": "2018-06-07T10:02:40.480Z",
          "content": "<p>No, I am not sure.  I think it should.</p>",
          "rawMarkdown": "No, I am not sure.  I think it should."
        }
      ]
    },
    {
      "id": 330067,
      "postDate": "2018-05-18T00:22:57.113Z",
      "content": "<p>.</p>",
      "rawMarkdown": ".\n"
    },
    {
      "id": 325529,
      "postDate": "2018-05-08T14:03:52.003Z",
      "content": "<p>other question about \"Not last\"-- is this mean when you do target mean,you ignore the last row of the group?</p>",
      "rawMarkdown": "other question about \"Not last\"-- is this mean when you do target mean,you ignore the last row of the group?",
      "replies": [
        {
          "id": 325547,
          "postDate": "2018-05-08T14:32:42.093Z",
          "content": "<p>not_last is not used in any grouping, it just helps the model learn that clicks happen preferably in the last row of each group.</p>",
          "rawMarkdown": "not_last is not used in any grouping, it just helps the model learn that clicks happen preferably in the last row of each group.",
          "votes": 1
        },
        {
          "id": 325676,
          "postDate": "2018-05-08T17:49:00.673Z",
          "content": "<p>Your question makes me think again. I added that feature quite late, but it could make sense to use it for grouping.  It probably would help.</p>",
          "rawMarkdown": "Your question makes me think again. I added that feature quite late, but it could make sense to use it for grouping.  It probably would help."
        },
        {
          "id": 327435,
          "postDate": "2018-05-11T14:54:40.960Z",
          "content": "<p>I think this is the key to make target encoding work (group by features like ['ip', 'os', 'device', 'not last']) :)\nsad I couldn't compete due to computer death :'(</p>",
          "rawMarkdown": "I think this is the key to make target encoding work (group by features like ['ip', 'os', 'device', 'not last']) :)\nsad I couldn't compete due to computer death :'(",
          "votes": 1
        }
      ]
    },
    {
      "id": 325523,
      "postDate": "2018-05-08T13:52:23.200Z",
      "content": "<p>Congrats @CPMP, a lot to learn from your solution.<br> </p>\n\n<pre><code>training on day &amp;lt;= 8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.\n</code></pre>\n\n<p>I had the same CV setting and hour 4 score was close to public LB,  but I failed miserably on private LB. Can I know how close is your validation score of other hours with private LB?</p>\n\n<pre><code>Features were computed on the concatenation of train and test_supplement\n</code></pre>\n\n<p>I was doing this in early days, but switching on making features for each day separately improved my validation scores + LB, so I  did not concat train + test_supplement anymore. May be that's why I fell on private LB. </p>",
      "rawMarkdown": "Congrats @CPMP, a lot to learn from your solution.<br> \n\n    training on day &lt;= 8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.\nI had the same CV setting and hour 4 score was close to public LB,  but I failed miserably on private LB. Can I know how close is your validation score of other hours with private LB?\n\n    Features were computed on the concatenation of train and test_supplement\nI was doing this in early days, but switching on making features for each day separately improved my validation scores + LB, so I  did not concat train + test_supplement anymore. May be that's why I fell on private LB. ",
      "replies": [
        {
          "id": 325550,
          "postDate": "2018-05-08T14:35:54.940Z",
          "content": "<p>Local score on the rest of test hours is not correlated with private LB as hour 4 is with public LB.  I mean, there is a larger gap, but they evolve in sync.  Any improvement on my local score translated in private LB improvement as far as I can see.</p>\n\n<p>You're right, you should have concatenated all data to compute your features.</p>",
          "rawMarkdown": "Local score on the rest of test hours is not correlated with private LB as hour 4 is with public LB.  I mean, there is a larger gap, but they evolve in sync.  Any improvement on my local score translated in private LB improvement as far as I can see.\n\nYou're right, you should have concatenated all data to compute your features.",
          "votes": 1
        },
        {
          "id": 325605,
          "postDate": "2018-05-08T16:12:23.423Z",
          "content": "<p>Right, but Ip encoding is different for each day, I don't get that how did you aggregate features like ip,app ratio of clicks per app? Did you aggregate them by day?<br> EDIT: Did you aggregated features like delta, frequency by day?</p>",
          "rawMarkdown": "Right, but Ip encoding is different for each day, I don't get that how did you aggregate features like ip,app ratio of clicks per app? Did you aggregate them by day?<br> EDIT: Did you aggregated features like delta, frequency by day?"
        }
      ]
    },
    {
      "id": 325360,
      "postDate": "2018-05-08T10:47:17.537Z",
      "content": "<p>Nice work! I notice that you just use 48 features. So how do you determine whether the feature you newly generate is useful or not? Thanks.</p>",
      "rawMarkdown": "Nice work! I notice that you just use 48 features. So how do you determine whether the feature you newly generate is useful or not? Thanks.",
      "replies": [
        {
          "id": 325418,
          "postDate": "2018-05-08T11:55:42.813Z",
          "content": "<p>Thanks.  I thought I explained it in my write up...  It is in the <strong>Validation</strong> section.  I basically add them one by one and keep what works.</p>",
          "rawMarkdown": "Thanks.  I thought I explained it in my write up...  It is in the **Validation** section.  I basically add them one by one and keep what works."
        }
      ]
    },
    {
      "id": 325343,
      "postDate": "2018-05-08T10:17:26.447Z",
      "content": "<p>I had many similar ideas as you used. Unfortunately I didn’t follow through hard enough and doubted many of them. Seeing you share them makes me feel validated and next time I will follow my intuitions with more faith.\nThanks!</p>",
      "rawMarkdown": "I had many similar ideas as you used. Unfortunately I didn’t follow through hard enough and doubted many of them. Seeing you share them makes me feel validated and next time I will follow my intuitions with more faith.\nThanks!",
      "replies": [
        {
          "id": 325420,
          "postDate": "2018-05-08T11:56:13.350Z",
          "content": "<p>Yes, you should try your ideas. That's the only way to know if they work or not.</p>",
          "rawMarkdown": "Yes, you should try your ideas. That's the only way to know if they work or not."
        }
      ]
    },
    {
      "id": 326185,
      "postDate": "2018-05-09T12:29:17.413Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 326734,
          "postDate": "2018-05-10T08:43:55.217Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 325511,
      "postDate": "2018-05-08T13:42:27.040Z",
      "content": "<p>Great job!\nThanks for sharing</p>",
      "rawMarkdown": "Great job!\nThanks for sharing",
      "votes": 1
    },
    {
      "id": 328609,
      "postDate": "2018-05-14T17:43:23.957Z",
      "content": "<p>Thanks for sharing your experience!</p>",
      "rawMarkdown": "Thanks for sharing your experience!",
      "votes": 2
    },
    {
      "id": 728433,
      "postDate": "2020-01-24T18:05:49.343Z",
      "content": "<p>Thank you very much for sharing !!</p>",
      "rawMarkdown": "Thank you very much for sharing !!"
    },
    {
      "id": 325281,
      "postDate": "2018-05-08T09:18:47.083Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 325525,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-05-08T13:54:14.053000",
      "content": "<p>Wow that's impressive. I love the matrix factorization part. Thanks for sharing in so much details (I wish I had access to your 250 messy files ....) Well done CPMP !</p>",
      "votes": 7,
      "replies": [
        {
          "id": 325548,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T14:33:32.937000",
          "content": "<p>Thanks Olivier, well done too on your side too.  I sincerely hope that the pair you form with Yifan will get the gold it deserves in a coming competition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 325572,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-05-08T15:06:32.453000",
          "content": "<p>Thanks for your kind words. My team mates were amazing and Yifan did an oustanding job. I would be hundreds of places behind without them! Looking at the data size on coming competitions I'm afraid I won't be able to join the party ;-)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 326842,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-10T12:20:35.643000",
      "content": "<p>I <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">shared my libFM implementation in Keras</a>.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 327712,
      "author_name": "Rohit Verma",
      "author_url": "",
      "post_date": "2018-05-12T08:45:23.070000",
      "content": "<p>Helpful indeed</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327834,
      "author_name": "Felix Houghton",
      "author_url": "",
      "post_date": "2018-05-12T15:41:56.977000",
      "content": "<p>Great writeup and well done for the 6th place finish!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 327721,
      "author_name": "RaviKaushik",
      "author_url": "",
      "post_date": "2018-05-12T09:15:36.173000",
      "content": "<p>very nice!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 326647,
      "author_name": "AmirH",
      "author_url": "",
      "post_date": "2018-05-10T05:14:33.733000",
      "content": "<p>@CPMP\nInteresting that you didn't use channel as categorical, in my first LGB model it was in top 3 predictors.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 326682,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T07:14:30.163000",
          "content": "<p>Hi, it was my top predictor according to lgb, but removing it did improve score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326750,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-05-10T09:14:34.917000",
          "content": "<p>@CPMP\nThat's really surprising,\nI thought of feature importance as a guide to which features I keep, I never imagined taking out a top predictor and see if it improves accuracy!</p>\n\n<p>How did you come up with the idea?\nHow did you test the first features' importance?\nDid you train on all fields and took away 1 feature at a time and see which helped accuracy or not?\n(run all without ip, run all without channel. etc...)</p>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326754,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T09:20:24.763000",
          "content": "<p>I ran quite a lot of experiments, maybe not 7,000 as Mamas did, but hundreds, probably around 500.  And removing original features was part of it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326623,
      "author_name": "HTan",
      "author_url": "",
      "post_date": "2018-05-10T04:04:28.647000",
      "content": "<p>Congrats and thanks to your sharing. I think I study a lot from your solution.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 326755,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T09:20:43.860000",
          "content": "<p>Thanks, glad you find it interesting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326423,
      "author_name": "Yair Beer",
      "author_url": "",
      "post_date": "2018-05-09T17:26:29.283000",
      "content": "<p>@CPMP thanks for sharing. I feel like a pleb next to your effort.\nI've noticed from a few mistakes I've done even from your few notes.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 326434,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T17:39:49.920000",
          "content": "<p>Thanks!  You did well too, I'd like to know how you get so much from ensembling given I basically failed at it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326506,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-05-09T20:57:05.667000",
          "content": "<p>I figured out why I could improve so much.\nBecause I didn't have enough RAM (52GB including swap), I trained the data only on 2 days (8,9 &amp; 7,9) and in different run using  14 of the hours with 26 features. Because I chose each time a different subset of the dataset the correlation wasn't perfect and I could improve so much.</p>\n\n<p>I used weighted mean with automatic anomaly removal. weight was related to LB score and correlation between the different submission. higher score -&gt; higher weight, lower correlation -&gt; higher weight but a bit less drastic.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326513,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T21:05:39.833000",
          "content": "<p>Thanks, makes sense now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326082,
      "author_name": "twoone",
      "author_url": "",
      "post_date": "2018-05-09T08:41:26.667000",
      "content": "<p>Hello，CPMP，thanks for your sharing。I am new to kaggle and I am not familiar with how to select features。When I construct a new feature ，how should I decide to add it or not？Just make a submission to see if it improves score？And when I construct some features，how should I decide they are good or not？Just make a submission to see if they improve score？Very thanks for your reply。</p>",
      "votes": 1,
      "replies": [
        {
          "id": 326086,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T08:50:24.900000",
          "content": "<blockquote>\n  <p>how should I decide to add it or not？</p>\n</blockquote>\n\n<p>Read again my writeup, I explain it.</p>\n\n<blockquote>\n  <p>Just make a submission to see if it improves score？</p>\n</blockquote>\n\n<p>Absolutely not.  That's what I called LB probing in y writeup.  This will make you feel good on public LB, and most probably result in a disaster on the private LB.  Key is to use local validation as I explain.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326062,
      "author_name": "Rayarrow",
      "author_url": "",
      "post_date": "2018-05-09T08:08:54.390000",
      "content": "<p>Congrats and Thanks! The MF part is really amazing!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 326112,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T09:57:19.193000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326023,
      "author_name": "Eric",
      "author_url": "",
      "post_date": "2018-05-09T07:03:50.680000",
      "content": "<p>By the way CPMP, how precious was intuitively the fact to use China day? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 326055,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T07:42:40.903000",
          "content": "<p>lag features helped but not much, maybe 0.0002.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326010,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "2018-05-09T06:52:07.540000",
      "content": "<p>Thanks for your sharing. Not finished reading your post yet!  But come here to comment with a big thumb up !  Very inspired about the \"Matrix factorization\" part !  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325979,
      "author_name": "Treebeard",
      "author_url": "",
      "post_date": "2018-05-09T06:32:13.207000",
      "content": "<p>Hi CPMP, Do you have any tips/resources to get better at feature engineering? This is the hardest part for any data scientist and something that i am not currently very good at. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 325989,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T06:40:24.477000",
          "content": "<ol>\n<li>Have a reliable CV setting.  Without it you cannot do FE.</li>\n<li>Try a lot, fail a lot, learn from failures.</li>\n<li>Read top teams writeup after each competition, look for things you did not think of, learn from them.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326016,
          "author_name": "Treebeard",
          "author_url": "",
          "post_date": "2018-05-09T06:59:16.123000",
          "content": "<p>Thanks for this, i will try to incorporate them religiously in the future. </p>\n\n<p>Are there any good resources as well to read up on FE? I usually read a few papers related to the topic and  try to embed some ideas from there. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326056,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T07:47:10.797000",
          "content": "<p>Here are two great presentations:</p>\n\n<p><a href=\"https://www.slideshare.net/OwenZhang2/tips-for-data-science-competitions\">https://www.slideshare.net/OwenZhang2/tips-for-data-science-competitions</a></p>\n\n<p><a href=\"https://www.slideshare.net/HJvanVeen/feature-engineering-72376750\">https://www.slideshare.net/HJvanVeen/feature-engineering-72376750</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326222,
          "author_name": "Treebeard",
          "author_url": "",
          "post_date": "2018-05-09T13:10:54.473000",
          "content": "<p>Thanks and congrats by the way :P </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325975,
      "author_name": "shark",
      "author_url": "",
      "post_date": "2018-05-09T06:23:40.737000",
      "content": "<p>Congrats CPMP and glad to see your writeup again.</p>\n\n<blockquote>\n  <p>China days. Introduced 24 periods that start at 4 pm. These were used\n  for lag features based on previous day(s) data.\n  Can you show some details and reasons about this feature ? </p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 325978,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T06:32:08.483000",
          "content": "<p>They make test look like the rest of the data.  They have about the same number of clicks per day.  UTC days are imbalanced, first and last UTC days on the data have way less clicks, esp day 6.  </p>\n\n<p>Because day 6 is so small, lag features computed for day 7 are not very good.  </p>\n\n<p>Using China day means the first 24 hours of data have no lag, but then all remaining data has lag computed form at least a full day.  This yield better result in my model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326009,
          "author_name": "shark",
          "author_url": "",
          "post_date": "2018-05-09T06:51:46.727000",
          "content": "<p>Thanks CPMP. It's a magic time handle.\nWill you share your code latter? Looking forward to seeing all your solutions especially matrix factorization.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 327439,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-11T14:58:41.103000",
          "content": "<p>HI I shared my matrix factorization code.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325871,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2018-05-09T01:48:17.173000",
      "content": "<p>Hello, </p>\n\n<p>Thanks for sharing your very intuitive solution in a very detailed manner. As a novice in this type of competitions I really appreciate it and see it as an opportunity to learn from others. I have questions on feature engineering and forgive me if these are stupid questions :)</p>\n\n<ul>\n<li><p><strong>delta with previous/next value</strong>, e.g. <code>time to next click by user</code>, when you are calculating this feature how do you create it for test/val data. Aren't we missing time until next click information for test set. <code>time since previous click</code> is easier to understand but if we have a long time horizon in test set wouldn't it introduce noise since we might be over exaggerating time. These made sense when I looked at Rossmann competition where time since/until promotions were used but I am confused when target is involved.</p></li>\n<li><p>Similarly for lagged <strong>previous target mean</strong>, are we using a sliding window to calculate the target encodings, e.g. calculate target encoding on previous X days(regularized) to encode next Y days (not regularized).  Let's say for this case let val set be 5 hours of period, use 1 hour sliding window to estimate features for next 5 hours, is this correct interpretation ?</p></li>\n</ul>\n\n<p>Thanks Again !                                                                                                                                                 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 325977,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T06:29:39.783000",
          "content": "<ol>\n<li><p>Features were computed on the concatenation of train and test_supplement.  test is a subset of test_supplement.  I did not compute features on test, I predict on test_supplement, then join.  I explained how I did it <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201\">here</a>.  Later, I added the not_last feature to the join columns.</p></li>\n<li><p>lagged features are computed on previous China day, not on a sliding window.</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326048,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2018-05-09T07:31:18.297000",
          "content": "<p>Thanks !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325739,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "2018-05-08T19:52:51.537000",
      "content": "<p>WOW,congrats and thanks for your writeup,it's really helpful.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325742,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T19:55:14.223000",
          "content": "<p>Thanks, means a lot coming from you.  I wish i had thought of RNN as you used them!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 325749,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2018-05-08T20:10:08.257000",
          "content": "<p>A competition is a journey full of uncertainty,we can not try too much in limited time,we must make decision everyday,even every hour,so if we can prove something,find something interesting or learn something,we are successful.<br>I always treat a competition as a force to learn more,I read many papers and solutions about CTR at the beginning of this competition,so I am looking foward to your keras-FM.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 326148,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-09T11:03:16.160000",
          "content": "<p>I will publish the libfm part soon.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326844,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T12:21:52.927000",
          "content": "<blockquote>\n  <p>I will publish the libfm part soon.</p>\n</blockquote>\n\n<p>Done, see <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325692,
      "author_name": "EV",
      "author_url": "",
      "post_date": "2018-05-08T18:13:16.287000",
      "content": "<p>CPMP, Congratulations! Thanks for sharing highlights of your solution. Your contributions to the Kaggle community are indeed remarkable and highly appreciated. The matrix factorization idea is ingenious. I agree with you that having a good CV methodology is key esp. for a competition with so much data. I also used three setups (small/medium/large) that successively validated my results. But I can see numerous other things I did not do and your post gives me so many ideas for the next set of competitions :-) So thanks again.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325744,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T19:55:47.830000",
          "content": "<p>Thanks a lot Adarsh!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325643,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2018-05-08T16:57:11.780000",
      "content": "<p>Thanks a lot for this detailed explanation. I was disappointed like a lot of other people to see my score dropped 400 place in 1h, and see people with 2 submissions beat me while I did hundreds of tries. </p>\n\n<p>But this was not the reason why I joined this competition, and it's definitely the high quality explanations / advises that you and others Kagglers provided that made me enjoyed the learning process. So thank you for that ! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325600,
      "author_name": "Araks Stepanyan",
      "author_url": "",
      "post_date": "2018-05-08T16:06:16.107000",
      "content": "<p>Thanks for sharing, I need to implement your feature selection strategies next time and practice getting rid of the features that took hours to engineer :)</p>\n\n<p>By the way, the ratio of the number of clicks per ip, day to the number of clicks per day was also working really well for me. Always in the top 15 features based on the feature importance. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325487,
      "author_name": "PW",
      "author_url": "",
      "post_date": "2018-05-08T13:16:26.200000",
      "content": "<p>Some question about the lag features:if the day is 7,the only data that you can use is data in the day 6,but the data amount in the day 6 is very small,this mean your lag feature in the day 7 is almost 0??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325502,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T13:35:28.583000",
          "content": "<p>Great question.  Lag features are undefined for first China day.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325456,
      "author_name": "Rui Li",
      "author_url": "",
      "post_date": "2018-05-08T12:41:27.470000",
      "content": "<p>great lesson learnt here is implementing matrix factorization for categorical combinations with large number of levels. I think I used most other methods similarly in this post, but didn't do the libfm, which makes the difference.  </p>\n\n<p>Definitely should learn that, and hopefully to use it in future competition and datasets.  Do you have any code base/implementation/method paper reference that we could look into? Thanks!!</p>\n\n<p>Again, thanks for great sharing and always something useful learnt here. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 325477,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T13:01:28.050000",
          "content": "<p>Start with <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html\">http://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html</a> and <a href=\"http://scikit-learn.org/stable/modules/decomposition.html\">http://scikit-learn.org/stable/modules/decomposition.html</a></p>\n\n<p>Truncated SVD is similar to PCA.</p>\n\n<p>For libfm, I'll write something on my blog on how to implement it in Keras.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 325489,
          "author_name": "Marcelo Senaga",
          "author_url": "",
          "post_date": "2018-05-08T13:17:50.460000",
          "content": "<p>What is the link to the blog? :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325501,
          "author_name": "Rui Li",
          "author_url": "",
          "post_date": "2018-05-08T13:33:54.437000",
          "content": "<p>@CPMP,  Thanks!! @Olecram, this is CPMP's blog,<a href=\"https://www.ibm.com/developerworks/community/blogs/jfp?lang=en\">https://www.ibm.com/developerworks/community/blogs/jfp?lang=en</a> \nanother good resource to learn. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 325522,
          "author_name": "Marcelo Senaga",
          "author_url": "",
          "post_date": "2018-05-08T13:50:59.437000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325450,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-05-08T12:31:23.623000",
      "content": "<p>Congratulations and thanks @CPMP for sharing your solution. </p>\n\n<p>I have learnt a few things from your FE i.e. the matrix factorization and the features you created to capture the leak. I now know my best single model with 21 features is nowhere near enough if you are talking 50 features :-)</p>\n\n<p>You are right about it being crazy to do this one solo. I decided to do it solo because this is my 1st pure time series competition hence; 1) I wanted to focus mainly on the learning part, 2) I felt like I do not have a lot of time hence may not be able to carry my weight to satisfaction if I joined a team. But no regrets as I have learnt so much.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325478,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T13:02:13.803000",
          "content": "<p>Bravo for you solo, sorry your rank was hurt by Dirk's late share.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325492,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2018-05-08T13:21:33.500000",
          "content": "<p>Thanks. Although I have only been doing this for roughly one year, I have learnt to expect such surprises in the last days of a competition. So I just shake it off and move on. The real surprise to me is that a blend of blend average ending up at #153 on the private LB with a silver medal.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325365,
      "author_name": "Laevatein",
      "author_url": "",
      "post_date": "2018-05-08T10:53:52.403000",
      "content": "<p>You really did a great job! Thx for sharing. You can have a good dream today. ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325361,
      "author_name": "antruong",
      "author_url": "",
      "post_date": "2018-05-08T10:47:50.360000",
      "content": "<p>Hi CPMP, thanks a lot for insightful sharing. You and top kagglers keep us up to the game as you guys showed that smarter ways can win the game, not just copy/blend others results. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325313,
      "author_name": "Rapexi",
      "author_url": "",
      "post_date": "2018-05-08T09:45:46.773000",
      "content": "<p>Well done CPMP</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325310,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-05-08T09:42:20.970000",
      "content": "<p>@CPMP Nice work!. I will repeat your results when i am free :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325421,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T11:57:40.117000",
          "content": "<p>Thanks!  Let me know if and when you start it, would be interested to see how it goes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325242,
      "author_name": "Eric",
      "author_url": "",
      "post_date": "2018-05-08T08:55:44.717000",
      "content": "<p>Hi CPMP, well done and thanks for this fruitful insights!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325263,
          "author_name": "Eric",
          "author_url": "",
          "post_date": "2018-05-08T09:06:55.833000",
          "content": "<p>I like your comment: my machine is a little bit slow! It is a <strong>20 core</strong> Xeon at 2.3 MHz with <strong>64 GB RAM</strong> and 64GB swap. This is way beyond my machine. And thanks for these nice details. Very interesting. Kaggle rocks and keeps me up at night!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 325269,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T09:13:02.863000",
          "content": "<p>Thanks!</p>\n\n<p>regarding our point on machines, any VM on a cloud is way faster.  A 4 cores i7 with SSD is faster.  A Macbook pro is faster.</p>\n\n<p>But yes, this is bigger and faster than an entry level laptop.  </p>\n\n<p>I used to compete on Kaggle with a macbook pro with 16GB.  I know how challenging it is to work with large datasets...  Net result is that I fried my macbookpro and decided to buy a relatively cheap machine that could scale out.  Yu can find second hand machines similar to it at a very reasonable price now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325289,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-05-08T09:24:49.847000",
          "content": "<p>For the 2nd half of the competition I was using a core i5 machine with 16GB RAM and 100+GB Swap space on 2 SSD's... I just shut down my machine after many weeks... Luckily I had an access to a better system in the first half and thats when I had quick and promising results..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 328627,
          "author_name": "Harjinder Singh",
          "author_url": "",
          "post_date": "2018-05-14T18:38:01.220000",
          "content": "<p>What OS  on machines, you folks are discussing? Linux/Window/MAC? I see only afordable Window machines for 32/64G in the market.  Linux very few and expensive.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 328673,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-14T21:29:56.677000",
          "content": "<p>You can always install Linux on a Windows machine.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337640,
          "author_name": "Rohit Verma",
          "author_url": "",
          "post_date": "2018-06-03T11:54:14.127000",
          "content": "<p>Yes , I have just installed kali</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325238,
      "author_name": "Samrat Pandiri",
      "author_url": "",
      "post_date": "2018-05-08T08:52:36.390000",
      "content": "<p>Thanks a ton for sharing this info <a href=\"/cpmpml\">@cpmpml</a> BTW do you plan to share any code, may be via github.. And thanks a lot for the mention tooo :-)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325241,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T08:55:21.057000",
          "content": "<p>I won't share the code in whole as it is messy, I didn't use source code control, rather copied files around. I have 250 files in my model directory...  Cleaning this and documenting is a daunting task, I am glad to not be in prize position for that ;)</p>\n\n<p>But I will share some of it via my blog.  I'll put pointer to it from here when I do it.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 325244,
          "author_name": "Eric",
          "author_url": "",
          "post_date": "2018-05-08T08:56:16.047000",
          "content": "<p>CPMP, where is exactly your blog?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325282,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T09:19:27.250000",
          "content": "<p>Here: <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp?lang=en\">https://www.ibm.com/developerworks/community/blogs/jfp?lang=en</a></p>\n\n<p>Not very active lately because of Kaggle.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 325284,
          "author_name": "Samrat Pandiri",
          "author_url": "",
          "post_date": "2018-05-08T09:21:05.710000",
          "content": "<p>Yeah snippets of code should do....</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326021,
          "author_name": "Eric",
          "author_url": "",
          "post_date": "2018-05-09T07:02:43.613000",
          "content": "<p>Thanks CPMP for the link about your blog. I will definitely keep an eye on it! And thank you in general for being so talkative on kaggle forum. You are my hero! No wonder why you are a discussion grand master and ranked n°1.. Keep it up. It is higly appreciated!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325235,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-05-08T08:47:05.573000",
      "content": "<p>Well done - a bucket full of awesome</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325445,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-08T12:25:58.697000",
          "content": "<p>You did well too, congrats!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 330163,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-18T07:05:52.493000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 330231,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-18T10:26:42.147000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 329923,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-17T15:05:31.850000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 329192,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-16T00:34:56.440000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 328884,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-15T09:08:33.700000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 328407,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-14T08:28:52.663000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327610,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-12T01:58:57.153000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327606,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-12T01:34:17.413000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327399,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-11T12:55:59.870000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327391,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-11T12:38:10.763000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327235,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-11T04:51:09.847000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 326672,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T06:43:08.703000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 326733,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T08:42:07.393000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326751,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T09:16:55.490000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326752,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T09:18:26.207000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326818,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T11:38:02.257000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326826,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T11:44:29.893000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 327028,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T17:42:23.867000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 327031,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T17:53:59.397000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326089,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-09T09:01:01.820000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 326093,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-09T09:04:20.823000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 326124,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-09T10:17:14.083000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326843,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T12:21:18.190000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 327258,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-11T06:34:26.533000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325484,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T13:11:40.677000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 325504,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T13:36:24.807000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 325517,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T13:47:15.443000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325520,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T13:49:13.390000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1316940,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-21T03:31:38.053000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 408114,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-22T10:38:36.190000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 333126,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-24T12:45:13.913000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 334239,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-26T18:50:32.083000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 339663,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-07T09:56:59.307000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 339666,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-07T10:02:40.480000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 330067,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-18T00:22:57.113000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325529,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T14:03:52.003000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 325547,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T14:32:42.093000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325676,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T17:49:00.673000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327435,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-11T14:54:40.960000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325523,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T13:52:23.200000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 325550,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T14:35:54.940000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325605,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T16:12:23.423000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325360,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T10:47:17.537000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 325418,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T11:55:42.813000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325343,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T10:17:26.447000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 325420,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-08T11:56:13.350000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326185,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-09T12:29:17.413000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 326734,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T08:43:55.217000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325511,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T13:42:27.040000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 328609,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-14T17:43:23.957000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 728433,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T18:05:49.343000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325281,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T09:18:47.083000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325218": "First of all, all those who managed to get decent submissions out of this huge dataset deserve kudos.  Even if you score is below Kirk's shared kernel.  What you did is way more valuable, and lessons learned here will help you later.  \n\nSecond, sharing is great when done in good faith, and lots of people did share a lot here, too many of them to name them.  Eve people who started here, like @Samrat, shared a lot.  This is what makes this community so valuable.\n\nThird, thanks to @inversion, Kaggle, and Talking data for organizing a very challenging, ans almost leak free competition.  I say almost because it is clear now that the test data was sorted by click time then target value.  Exploiting this was the final twist that helped some of us fare better.  But the impact is not that large, I estimate it to be about 0.0004 for me.  And it was [disclosed soon enough][1] for everyone to react to it.  Thanks to @plantsgo for sharing it soon enough.\n\nI didn't decide to go solo from the start, but as time went by, I saw I was making progress every day, and decided to go solo till the end.  In retrospect I am not sure it was wise, I didn't sleep much in the last week ;)  I did receive some invites to merge during the last week before deadline, and I thank people for them.  I did miss a very late invite from a top 10 team as I was away that evening.  I wonder what would have been our score if we had teamed.\n\nAnyway here is my solution.  Given I was solo, and given the size of the data set, which meant hours to produce a submission, I decided to focus. I focused on a single type of model, LightGBM.  I split my time roughly as follows:\n\n - 80% feature engineering\n - 10% making local validation as fast as possible\n - 5% hyper parameter tuning\n - 5% ensembling\n\nI spent most of my time doing feature selection, as my machine was not usable with 50 features or more.  I wish I had used the [trick shared by Kruegger][2], maybe I would have been able to add more features.  However, being forced to be selective about features probably led to better models in the end.  And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by @authman.\n\nMy private LB score comes from a single lgb run with 48 features that scored 0.9825 public and 0.9835 private.  I submitted a blend of this with 5 other similar models that yield 0.9828 public and 0.9837 private, but for some reason [that sub isn't taken into account][3].  This is OK as my rank would not change with it.  I didn't submit all the 5 models individually during the competition, but did it after to get their score.  The best single lgb run scores 0.9827 public and 0.9836 private, with 48 features.\n\nI mostly used a 20 core Xeon at 2.3 MHz with 64 GB RAM and 64GB swap.  That machine is a bit slow, but it scales when using 20 threads.  I also used another machine with a 4 core i7 and 2 GPU to run Keras (see below).\n\n**Validation**\n\nLet's us look at validation.  This is key.  If you don't have a good validation scheme then you rely solely on LB probing, which can easily lead to overfit.  I ended up settling on:\n\n - training on day &lt;=  8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.\n - retraining on all data using 1.2 times the number of trees found by early stopping in validation\n\nUsing two validation sets was to make sure I was not overfiting to one of them.  The hours were selected to match the public and private test data.  I also watched the train auc in the early days, discarding features that improved validation but also increased the gap with train a lot. I stopped watching train auc in the last week to speed up things, but last time I checked I had a quite small gap.\n\nI also used LB, i.e. only kept a feature if local auc and LB improved. Yes, I know this can lead to overfit, but given I was filtering first on local validation I think I escaped it for the most part.\n\n This was a very effective scheme, with the hour 4 score being the same as LB score with a std difference around 0.0001.  However, it is very time consuming because of the computation of the auc metric for early stopping.  In order to speed it for feature evaluation I used two lighter ways.  First, using only day 9 data, with 5% of hour 4 data for validation, the rest for training.  This could run in less than one hour, and was used as a filter.  Only features that improved on that went to the next stage which was train on day 8 and validate on day 9, both hour 4 and other test hours.  This was also a very effective scheme, with very good correlation with LB score, but it was not effective when evaluating lag features.  I therefore switched to training on day &lt;= 8 later on.\n\nAnother way of speeding feature evaluation was to share each feature in a separate feather file.  This way, testing a feature set only requires assembling a set of files into one dataset.  Features were mostly tested by adding them one by one, and keeping them if local validation score improved by at least 0.00005.  I also added several of them at once, then removed them one by one to see if validation score decreased.  I basically did feature selection full time for the competition, preparing experiments to be run while I was away during day, or while I was sleeping.  The machine never stopped.\n\n**Feature Engineering**\n\nFeatures were computed on the concatenation of train and test_supplement, sorted by click time then by original order.  Now I am not sure the second item was useful.\n\nI used several families of features.\n\n -  Only  app, and os from the original features were kept. They were handled as categorical, and were my strongest 2 features with a third category made of the hour in the day.\n - China days.  Introduced 24 periods that start at 4 pm.  These were\n   used for lag features based on previous day(s) data.\n - User: ip, device, os triplets.  \n - Aggregates on various feature groups, similar to what was shared in many public kernels.  Aggregates I used were count, count of unique values, delta with previous value, delta with next value.  Time to next click when grouped by user was important.  Other useful ones I didn't see in kernels: delta with previous app.\n - Lag features, based on previous China days values.  Previous count by some grouping, and previous target mean by some grouping.  The latter was a weighted average with the overall target mean, the weights being such that groups with few rows in it had a value closer to the overall average.  This is a standard normalization in target encoding.\n - Ratios like number of clicks per ip, app to number of click per app.\n - Not last.  This was to capture the leak.  It is one except for rows that are not the last of their group when grouped by user, app, and click time.  I ignored channel as I think that clicks are attributed to the most recent click having same user and app as the download.\n - Target.  This is to also capture the leak.  I modified the target in train data by sorting is_attributed within group by user, app, and click time. The combination of both ways to capture the leak led to a boost between 0.0004 and 0.0005.\n - Matrix factorization.  This was to capture the similarity between users and app.  I use several of them.  They all start with the same approach; construct a matrix with log of click counts. I used: ip x app, user x app, and os x device x app.  These matrices are extremely sparse (most values are 0).  For the first two I used truncated svd from sklearn, which gives me latent vectors (embeddings) for ip and user.  For the last one, given there are 3 factors, I implemented libfm in Keras and used the embeddings it computes.  I used between 3 and 5 latent factors.  All in all, these embeddings gave me a boost over 0.0010.  I think this is what led me in top 10.  I got some variety of models by varying which embeddings I was using.\n\n**Hyper parameter tuning**\n\nI spent time given how long it is to run an experiment, but I didn't tune much.  Main settings were to scale positive by around 400, use an initial score that minimizes expected loss if target is constant, and min child per leaf to be such that it requires at least 1 positive plus another example, in order to avoid overfiting to single positive examples.  I used 31 leaves and a depth of 8.  \n\n**Ensembling**\n\nOn my local validation, the best way to blend several models was to average the logit of the predictions (aka raw predictions).  I started doing restacking, i.e. adding validation predictions to day 9 features, and training on it, but this was hitting my 50 feature limit, and runs were very long.  I did not ran it the last day for that reason.  It may have given me a little additional boost, but I don't think it would have been enough to move up in the LB, because my models were not diverse enough anyway.  I also think that using only day 9 for second level was leaving too much on the table.  I thought of generating of prediction for the full dataset, but that was a daunting task.  I now see this is what @bestfitting did, great move on his part.\n\n**Takeaway**\n\nI think it was crazy to do this solo, too much work for a single person.  I really admire the other fools that went same way.  Once solo, I am not sure I should have done things differently, except for spending time to alleviate the 50 features limit.\n\nAlso, as often in my competitions, I make a lot of progress the last day, not sure why.  In this case I moved from 0.9824 to 0.9828 public, and 0.9832 to 0.9837 private.  The lesson is to never give up, and not let the public LB dictate your mood.\n\nI hope the above will be useful to some.  Thanks for reading it all ;)\n\nEdit: I [shared my libFM implementation][4].\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56105\n  [3]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56234\n  [4]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497",
    "325525": "Wow that's impressive. I love the matrix factorization part. Thanks for sharing in so much details (I wish I had access to your 250 messy files ....) Well done CPMP !",
    "326842": "I [shared my libFM implementation in Keras][1].\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497",
    "327712": "Helpful indeed",
    "327834": "Great writeup and well done for the 6th place finish!",
    "327721": "very nice!!",
    "326647": "@CPMP\nInteresting that you didn't use channel as categorical, in my first LGB model it was in top 3 predictors.",
    "326623": "Congrats and thanks to your sharing. I think I study a lot from your solution.",
    "326423": "@CPMP thanks for sharing. I feel like a pleb next to your effort.\nI've noticed from a few mistakes I've done even from your few notes.\n",
    "326082": "Hello，CPMP，thanks for your sharing。I am new to kaggle and I am not familiar with how to select features。When I construct a new feature ，how should I decide to add it or not？Just make a submission to see if it improves score？And when I construct some features，how should I decide they are good or not？Just make a submission to see if they improve score？Very thanks for your reply。",
    "326062": "Congrats and Thanks! The MF part is really amazing!!!",
    "326023": "By the way CPMP, how precious was intuitively the fact to use China day? ",
    "326010": "Thanks for your sharing. Not finished reading your post yet!  But come here to comment with a big thumb up !  Very inspired about the \"Matrix factorization\" part !  ",
    "325979": "Hi CPMP, Do you have any tips/resources to get better at feature engineering? This is the hardest part for any data scientist and something that i am not currently very good at. ",
    "325975": "Congrats CPMP and glad to see your writeup again.\n\n&gt; China days. Introduced 24 periods that start at 4 pm. These were used\n&gt; for lag features based on previous day(s) data.\nCan you show some details and reasons about this feature ? \n",
    "325871": "Hello, \n\nThanks for sharing your very intuitive solution in a very detailed manner. As a novice in this type of competitions I really appreciate it and see it as an opportunity to learn from others. I have questions on feature engineering and forgive me if these are stupid questions :)\n\n-  **delta with previous/next value**, e.g. `time to next click by user`, when you are calculating this feature how do you create it for test/val data. Aren't we missing time until next click information for test set. `time since previous click` is easier to understand but if we have a long time horizon in test set wouldn't it introduce noise since we might be over exaggerating time. These made sense when I looked at Rossmann competition where time since/until promotions were used but I am confused when target is involved.\n\n- Similarly for lagged **previous target mean**, are we using a sliding window to calculate the target encodings, e.g. calculate target encoding on previous X days(regularized) to encode next Y days (not regularized).  Let's say for this case let val set be 5 hours of period, use 1 hour sliding window to estimate features for next 5 hours, is this correct interpretation ?\n\nThanks Again !                                                                                                                                                 ",
    "325739": "WOW,congrats and thanks for your writeup,it's really helpful.",
    "325692": "CPMP, Congratulations! Thanks for sharing highlights of your solution. Your contributions to the Kaggle community are indeed remarkable and highly appreciated. The matrix factorization idea is ingenious. I agree with you that having a good CV methodology is key esp. for a competition with so much data. I also used three setups (small/medium/large) that successively validated my results. But I can see numerous other things I did not do and your post gives me so many ideas for the next set of competitions :-) So thanks again.",
    "325643": "Thanks a lot for this detailed explanation. I was disappointed like a lot of other people to see my score dropped 400 place in 1h, and see people with 2 submissions beat me while I did hundreds of tries. \n\nBut this was not the reason why I joined this competition, and it's definitely the high quality explanations / advises that you and others Kagglers provided that made me enjoyed the learning process. So thank you for that ! ",
    "325600": "Thanks for sharing, I need to implement your feature selection strategies next time and practice getting rid of the features that took hours to engineer :)\n\nBy the way, the ratio of the number of clicks per ip, day to the number of clicks per day was also working really well for me. Always in the top 15 features based on the feature importance. ",
    "325487": "Some question about the lag features:if the day is 7,the only data that you can use is data in the day 6,but the data amount in the day 6 is very small,this mean your lag feature in the day 7 is almost 0??",
    "325456": "great lesson learnt here is implementing matrix factorization for categorical combinations with large number of levels. I think I used most other methods similarly in this post, but didn't do the libfm, which makes the difference.  \n\nDefinitely should learn that, and hopefully to use it in future competition and datasets.  Do you have any code base/implementation/method paper reference that we could look into? Thanks!!\n\nAgain, thanks for great sharing and always something useful learnt here. ",
    "325450": "Congratulations and thanks @CPMP for sharing your solution. \n\nI have learnt a few things from your FE i.e. the matrix factorization and the features you created to capture the leak. I now know my best single model with 21 features is nowhere near enough if you are talking 50 features :-)\n\nYou are right about it being crazy to do this one solo. I decided to do it solo because this is my 1st pure time series competition hence; 1) I wanted to focus mainly on the learning part, 2) I felt like I do not have a lot of time hence may not be able to carry my weight to satisfaction if I joined a team. But no regrets as I have learnt so much.",
    "325365": "You really did a great job! Thx for sharing. You can have a good dream today. ;)",
    "325361": "Hi CPMP, thanks a lot for insightful sharing. You and top kagglers keep us up to the game as you guys showed that smarter ways can win the game, not just copy/blend others results. ",
    "325313": "Well done CPMP",
    "325310": "@CPMP Nice work!. I will repeat your results when i am free :)",
    "325242": "Hi CPMP, well done and thanks for this fruitful insights!",
    "325238": "Thanks a ton for sharing this info @cpmpml BTW do you plan to share any code, may be via github.. And thanks a lot for the mention tooo :-)",
    "325235": "Well done - a bucket full of awesome",
    "330163": "Thank you very much for sharing the detailed information, @cpmpml.\n\n&gt; And I would not have been able to move past40 features without using the 'two_round' parameter as suggested by @autheman.\n\nWhat are past40 features and 'two_round' parameter?\n",
    "329923": "Thanks for sharing! Very interesting for matrix factorization part!",
    "329192": "That very helpful",
    "328884": "very helpful!",
    "328407": "really helpful!",
    "327610": "so cool",
    "327606": "Great work",
    "327399": "good",
    "327391": "Great work@CPMP! Congratulations and thanks for detailed sharing, your tips are really useful!",
    "327235": "Hello all. I'm brand new around here, having just completed the 'my first model' training in the learn section. Look forward to the day when I can leave substantive comments because I understand what y'all are talking about. :)",
    "326672": "Thank your for your sharing @CPMP. I also did matrix factorization, using ip x app (my machine can't handle the size of user x app) with count of clicks instead of log counts you used. My procedure is to use ALS to extract 4 or 8 factors corresponding to ip, however both set of factors doesn't improve my validation score and I'm wondering whether it is a problem of my way of counting, ALS, or using only ip x app.  Please share your thoughts with me.",
    "326089": "Thanks a lot fortaking the time to explain and share your solutions and insight ! I'm not sure I fully understand the matrix factorization part, which seems really interesting. Do you have any detailed explanation of this ? (article, code, anything really ?)",
    "325484": "Very detailed description, I read it carefully and it contains many golden experiences. I must say thank you for your endevor. \n\nBTW, there's one thing I would like to know a bit more: you mentioned that you used 1.2 times num_boost_round found by early stopping on validation set. How is 1.2 determined? Is that determined by the ratio of samples on training set for local validation over training set for submission?",
    "1316940": "Hi @cpmpml \n\nI saw here you were using a slow computer for your work, would your results or work process be better if you had more GPU's ?",
    "408114": "Thank you for sharing this Sir.It helped me a lot....",
    "333126": " \n\n&gt; and min child per leaf to be such that it requires at least 1 positive plus another example\n\nHow did you achieve that @cpmpml ? parameter value(which) or modifying source code? Whould you mind sharing?\nThanks in advance and congrats",
    "330067": ".\n",
    "325529": "other question about \"Not last\"-- is this mean when you do target mean,you ignore the last row of the group?",
    "325523": "Congrats @CPMP, a lot to learn from your solution.<br> \n\n    training on day &lt;= 8, and validating on both day 9 - hour 4, and day-9, hours 5, 9, 10, 13, 14.\nI had the same CV setting and hour 4 score was close to public LB,  but I failed miserably on private LB. Can I know how close is your validation score of other hours with private LB?\n\n    Features were computed on the concatenation of train and test_supplement\nI was doing this in early days, but switching on making features for each day separately improved my validation scores + LB, so I  did not concat train + test_supplement anymore. May be that's why I fell on private LB. ",
    "325360": "Nice work! I notice that you just use 48 features. So how do you determine whether the feature you newly generate is useful or not? Thanks.",
    "325343": "I had many similar ideas as you used. Unfortunately I didn’t follow through hard enough and doubted many of them. Seeing you share them makes me feel validated and next time I will follow my intuitions with more faith.\nThanks!",
    "326185": "",
    "325511": "Great job!\nThanks for sharing",
    "328609": "Thanks for sharing your experience!",
    "728433": "Thank you very much for sharing !!",
    "325281": "Thanks for sharing."
  }
}