{
  "id": 56475,
  "title": "1st place solution",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56475",
  "author_name": "Komaki",
  "post_date": "2018-05-10T07:08:37.897000",
  "votes": 270,
  "comment_count": 95,
  "views": 0,
  "content": "<p>We'd like to thank TalkingData and Kaggle for organizing this exciting competition. This competition gave us a fantastic opportunity to learn how to deal with very large table data. <br>\n<br>\nHere is our solution. <br>\n<br>\n<strong>Strategy:</strong> <br>\nOur solution heavily depends on negative down-sampling [1, 2], which means we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples, but we didn't see much performance deterioration when we tested with our initial features. Moreover, we could get better performance when creating a submission by bagging five predictors trained on five sampled datasets created from different random seeds. This technique allowed us to use hundreds of features while keeping LGB training time less than 30 minutes. <br>\n<br>\n<strong>Features:</strong> <br>\nFirst, we started from features from Kernels. On Thanks pranav84 and other Kagglers for sharing their awesome insights! We did feature engineerings using all the data examples instead of the down-sampled ones. <br>\n- five raw categorical features (ip, os, app, channel, device) <br>\n- time categorical features (day, hour) <br>\n- some count features <br>\nThen, we created a bunch of features in a brute-force way. For each combination of five raw categorical features (ip, os, app, channel, and device), we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features): <br>\n- click count within next one/six hours <br>\n- forward/backward click time delta <br>\n- average attributed ratio of past clicks <br>\nWe didn't do feature selection. We just added all of them to our model. At that point, our LGB model's score was 0.9808. <br>\n<br>\nNext, we tried categorical feature embedding by using LDA/NMF/LSA. Here is the pseudo code to compute LDA topics of IPs related to app. (LDA is latent Dirichlet allocation)<br></p>\n\n<pre>apps_of_ip = {}\nfor sample in data_samples:\n  apps_of_ip.setdefault(sample['ip'], []).append(str(sample['app']))\nips = list(apps_of_ip.keys())\napps_as_sentence = [' '.join(apps_of_ip[ip]) for ip in ips]\napps_as_matrix = CountTokenizer().fit_transform(apps_as_sentence)\ntopics_of_ips = LDA(n_components=5).fit_transform(apps_as_matrix)\n</pre>\n\n<p>We computed this feature for all the 20 (=5*(5-1)) combinations of 5 raw features and set the topic size to 5. This ended up with 100 new features. We also computed similar features using NMF and PCA, in total 300 new features. 0.9821 with a single LGB. <br>\n<br>\nAfter that, we removed all raw categorical features except app since we supposed embedding features cover information available from them. Surprisingly, this minor change made our public LB score jump up from 0.9821 to 0.9828. Actually, we don't know what causes this significant score improvement. <br>\nBesides features mentioned here, we created higher dimensional LDA features and features that try to address the duplicate sample problem. These features somewhat improve our public LB score. <br>\n<br>\n<strong>Models:</strong> <br>\nWe used day 7 &amp; 8 for training and day 9 for validation, and chose the best number of iterations of LGB. Then, we trained a model on day 7 &amp; 8 &amp; 9 with the obtained number of iterations for creating submission. After we finished feature engineering, flowlight's five-bagged LGB model reached 0.98333 on public LB (and 0.98420 on private LB), which was trained on 646 features. As far as we remember, the memory usage for training this model was less than 100GB (&lt;64GB will be possible with minor code modification). <br>\nI implemented a simple three layer NN model as some kernels do. It scored worse than LGB models by 0.0013 points with 0.005 down-sampling rate at first. Then, I realized it should be trained with more negative data samples, probably. However, we didn't afford to use so many examples because of the massive amount of our features (and my bad implementation, disk space, no GPU, close deadline, etc...). The final three-bagged NN model scored 0.98258 on public LB. <br>\nWe made our final submission with a rank-based weighted averaging. It is composed of seven bagged LGB models and a single bagged NN  It scored 0.98343 on public LB. <br>\n<br>\n<br>\n[1] <a href=\"http://quinonero.net/Publications/predicting-clicks-facebook.pdf\">Practical Lessons from Predicting Clicks on Ads at\nFacebook</a> <br>\n[2] <a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\">Ad Click Prediction: a View from the Trenches</a> <br>\n<br>\nEDIT: Added some extra explanations based on frequently asked questions in comments.</p>",
  "messages": [
    {
      "id": 326680,
      "postDate": "2018-05-10T07:08:37.897Z",
      "content": "<p>We'd like to thank TalkingData and Kaggle for organizing this exciting competition. This competition gave us a fantastic opportunity to learn how to deal with very large table data. <br>\n<br>\nHere is our solution. <br>\n<br>\n<strong>Strategy:</strong> <br>\nOur solution heavily depends on negative down-sampling [1, 2], which means we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples, but we didn't see much performance deterioration when we tested with our initial features. Moreover, we could get better performance when creating a submission by bagging five predictors trained on five sampled datasets created from different random seeds. This technique allowed us to use hundreds of features while keeping LGB training time less than 30 minutes. <br>\n<br>\n<strong>Features:</strong> <br>\nFirst, we started from features from Kernels. On Thanks pranav84 and other Kagglers for sharing their awesome insights! We did feature engineerings using all the data examples instead of the down-sampled ones. <br>\n- five raw categorical features (ip, os, app, channel, device) <br>\n- time categorical features (day, hour) <br>\n- some count features <br>\nThen, we created a bunch of features in a brute-force way. For each combination of five raw categorical features (ip, os, app, channel, and device), we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features): <br>\n- click count within next one/six hours <br>\n- forward/backward click time delta <br>\n- average attributed ratio of past clicks <br>\nWe didn't do feature selection. We just added all of them to our model. At that point, our LGB model's score was 0.9808. <br>\n<br>\nNext, we tried categorical feature embedding by using LDA/NMF/LSA. Here is the pseudo code to compute LDA topics of IPs related to app. (LDA is latent Dirichlet allocation)<br></p>\n\n<pre>apps_of_ip = {}\nfor sample in data_samples:\n  apps_of_ip.setdefault(sample['ip'], []).append(str(sample['app']))\nips = list(apps_of_ip.keys())\napps_as_sentence = [' '.join(apps_of_ip[ip]) for ip in ips]\napps_as_matrix = CountTokenizer().fit_transform(apps_as_sentence)\ntopics_of_ips = LDA(n_components=5).fit_transform(apps_as_matrix)\n</pre>\n\n<p>We computed this feature for all the 20 (=5*(5-1)) combinations of 5 raw features and set the topic size to 5. This ended up with 100 new features. We also computed similar features using NMF and PCA, in total 300 new features. 0.9821 with a single LGB. <br>\n<br>\nAfter that, we removed all raw categorical features except app since we supposed embedding features cover information available from them. Surprisingly, this minor change made our public LB score jump up from 0.9821 to 0.9828. Actually, we don't know what causes this significant score improvement. <br>\nBesides features mentioned here, we created higher dimensional LDA features and features that try to address the duplicate sample problem. These features somewhat improve our public LB score. <br>\n<br>\n<strong>Models:</strong> <br>\nWe used day 7 &amp; 8 for training and day 9 for validation, and chose the best number of iterations of LGB. Then, we trained a model on day 7 &amp; 8 &amp; 9 with the obtained number of iterations for creating submission. After we finished feature engineering, flowlight's five-bagged LGB model reached 0.98333 on public LB (and 0.98420 on private LB), which was trained on 646 features. As far as we remember, the memory usage for training this model was less than 100GB (&lt;64GB will be possible with minor code modification). <br>\nI implemented a simple three layer NN model as some kernels do. It scored worse than LGB models by 0.0013 points with 0.005 down-sampling rate at first. Then, I realized it should be trained with more negative data samples, probably. However, we didn't afford to use so many examples because of the massive amount of our features (and my bad implementation, disk space, no GPU, close deadline, etc...). The final three-bagged NN model scored 0.98258 on public LB. <br>\nWe made our final submission with a rank-based weighted averaging. It is composed of seven bagged LGB models and a single bagged NN  It scored 0.98343 on public LB. <br>\n<br>\n<br>\n[1] <a href=\"http://quinonero.net/Publications/predicting-clicks-facebook.pdf\">Practical Lessons from Predicting Clicks on Ads at\nFacebook</a> <br>\n[2] <a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\">Ad Click Prediction: a View from the Trenches</a> <br>\n<br>\nEDIT: Added some extra explanations based on frequently asked questions in comments.</p>",
      "rawMarkdown": "We'd like to thank TalkingData and Kaggle for organizing this exciting competition. This competition gave us a fantastic opportunity to learn how to deal with very large table data. <br>\n<br>\nHere is our solution. <br>\n<br>\n**Strategy:** <br>\nOur solution heavily depends on negative down-sampling [1, 2], which means we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples, but we didn't see much performance deterioration when we tested with our initial features. Moreover, we could get better performance when creating a submission by bagging five predictors trained on five sampled datasets created from different random seeds. This technique allowed us to use hundreds of features while keeping LGB training time less than 30 minutes. <br>\n<br>\n**Features:** <br>\nFirst, we started from features from Kernels. On Thanks pranav84 and other Kagglers for sharing their awesome insights! We did feature engineerings using all the data examples instead of the down-sampled ones. <br>\n- five raw categorical features (ip, os, app, channel, device) <br>\n- time categorical features (day, hour) <br>\n- some count features <br>\nThen, we created a bunch of features in a brute-force way. For each combination of five raw categorical features (ip, os, app, channel, and device), we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features): <br>\n- click count within next one/six hours <br>\n- forward/backward click time delta <br>\n- average attributed ratio of past clicks <br>\nWe didn't do feature selection. We just added all of them to our model. At that point, our LGB model's score was 0.9808. <br>\n<br>\nNext, we tried categorical feature embedding by using LDA/NMF/LSA. Here is the pseudo code to compute LDA topics of IPs related to app. (LDA is latent Dirichlet allocation)<br>\n<pre>apps_of_ip = {}\nfor sample in data_samples:\n  apps_of_ip.setdefault(sample['ip'], []).append(str(sample['app']))\nips = list(apps_of_ip.keys())\napps_as_sentence = [' '.join(apps_of_ip[ip]) for ip in ips]\napps_as_matrix = CountTokenizer().fit_transform(apps_as_sentence)\ntopics_of_ips = LDA(n_components=5).fit_transform(apps_as_matrix)\n</pre>\nWe computed this feature for all the 20 (=5*(5-1)) combinations of 5 raw features and set the topic size to 5. This ended up with 100 new features. We also computed similar features using NMF and PCA, in total 300 new features. 0.9821 with a single LGB. <br>\n<br>\nAfter that, we removed all raw categorical features except app since we supposed embedding features cover information available from them. Surprisingly, this minor change made our public LB score jump up from 0.9821 to 0.9828. Actually, we don't know what causes this significant score improvement. <br>\nBesides features mentioned here, we created higher dimensional LDA features and features that try to address the duplicate sample problem. These features somewhat improve our public LB score. <br>\n<br>\n**Models:** <br>\nWe used day 7 &amp; 8 for training and day 9 for validation, and chose the best number of iterations of LGB. Then, we trained a model on day 7 &amp; 8 &amp; 9 with the obtained number of iterations for creating submission. After we finished feature engineering, flowlight's five-bagged LGB model reached 0.98333 on public LB (and 0.98420 on private LB), which was trained on 646 features. As far as we remember, the memory usage for training this model was less than 100GB (&lt;64GB will be possible with minor code modification). <br>\nI implemented a simple three layer NN model as some kernels do. It scored worse than LGB models by 0.0013 points with 0.005 down-sampling rate at first. Then, I realized it should be trained with more negative data samples, probably. However, we didn't afford to use so many examples because of the massive amount of our features (and my bad implementation, disk space, no GPU, close deadline, etc...). The final three-bagged NN model scored 0.98258 on public LB. <br>\nWe made our final submission with a rank-based weighted averaging. It is composed of seven bagged LGB models and a single bagged NN  It scored 0.98343 on public LB. <br>\n<br>\n<br>\n[1] <a href=\"http://quinonero.net/Publications/predicting-clicks-facebook.pdf\">Practical Lessons from Predicting Clicks on Ads at\nFacebook</a> <br>\n[2] <a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\">Ad Click Prediction: a View from the Trenches</a> <br>\n<br>\nEDIT: Added some extra explanations based on frequently asked questions in comments.",
      "votes": 270
    },
    {
      "id": 326747,
      "postDate": "2018-05-10T09:09:32.277Z",
      "content": "<p>Congrats,a very good solution,your 1st place is well deserved,thanks for sharing. I think if we can combine the top teams' solutions ,we will get a very high score.</p>",
      "rawMarkdown": "Congrats,a very good solution,your 1st place is well deserved,thanks for sharing. I think if we can combine the top teams' solutions ,we will get a very high score.",
      "votes": 9,
      "replies": [
        {
          "id": 326757,
          "postDate": "2018-05-10T09:23:19.427Z",
          "content": "<p>I am not sure I qualify as one of the top teams, but I shared my 0.98379 private solution <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423\">here</a>.  Maybe others could do the same so that we can blend?</p>",
          "rawMarkdown": "I am not sure I qualify as one of the top teams, but I shared my 0.98379 private solution [here][1].  Maybe others could do the same so that we can blend?\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423",
          "votes": 4
        },
        {
          "id": 326761,
          "postDate": "2018-05-10T09:30:09.453Z",
          "content": "<p>Absolutely should include you  :)<br>To get top 50 is not so easy in this competition.<br></p>",
          "rawMarkdown": "Absolutely should include you  :)<br>To get top 50 is not so easy in this competition.<br>\n",
          "votes": 3
        },
        {
          "id": 327078,
          "postDate": "2018-05-10T19:31:14.327Z",
          "content": "<p>I feel the same. I am very impressed to see so many wonderful ideas that I have never thought of. </p>",
          "rawMarkdown": "I feel the same. I am very impressed to see so many wonderful ideas that I have never thought of. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 326793,
      "postDate": "2018-05-10T10:55:21.323Z",
      "content": "<p>Congratulations <a href=\"/ckomaki\">@ckomaki</a> and <a href=\"/flowlight\">@flowlight</a>. Excellent approach and very well deserved win. And a big thank you for the mention :) This actually made my day! </p>",
      "rawMarkdown": "Congratulations @ckomaki and @flowlight. Excellent approach and very well deserved win. And a big thank you for the mention :) This actually made my day! ",
      "votes": 8,
      "replies": [
        {
          "id": 326932,
          "postDate": "2018-05-10T14:27:23.777Z",
          "content": "<p>@Pranav, It is a pity to see how Dirk share impacted you.  Most people reused your shared kernels.</p>",
          "rawMarkdown": "@Pranav, It is a pity to see how Dirk share impacted you.  Most people reused your shared kernels.",
          "votes": 1
        },
        {
          "id": 326946,
          "postDate": "2018-05-10T14:54:21.077Z",
          "content": "<p>That's alrite @CPMP. Though Dirk impacted me too but I also blame myself for making stupid selections on judgement day :) But honestly, this competition was really a great learning experience for me specifically in feature engineering part. </p>\n\n<p>Shared learning from community actually helped me a lot and I'm sure that this will contribute positively for me in next competitions :) </p>",
          "rawMarkdown": "That's alrite @CPMP. Though Dirk impacted me too but I also blame myself for making stupid selections on judgement day :) But honestly, this competition was really a great learning experience for me specifically in feature engineering part. \n\nShared learning from community actually helped me a lot and I'm sure that this will contribute positively for me in next competitions :) ",
          "votes": 7
        }
      ]
    },
    {
      "id": 326714,
      "postDate": "2018-05-10T08:07:04.547Z",
      "content": "<p>fantastic solution, thank you very much for sharing - Matrix factorization for feature interaction will be a main take away for me from this competition </p>",
      "rawMarkdown": "fantastic solution, thank you very much for sharing - Matrix factorization for feature interaction will be a main take away for me from this competition ",
      "votes": 3
    },
    {
      "id": 326705,
      "postDate": "2018-05-10T07:49:28.920Z",
      "content": "<p>I must say that I feel stupid for not having tested downsampling of negative.  I usually try it on very imbalanced datasets.  Well done!</p>\n\n<p>Are you computing your lsa models on downsampled data or on the original one?</p>",
      "rawMarkdown": "I must say that I feel stupid for not having tested downsampling of negative.  I usually try it on very imbalanced datasets.  Well done!\n\nAre you computing your lsa models on downsampled data or on the original one?",
      "votes": 4,
      "replies": [
        {
          "id": 326729,
          "postDate": "2018-05-10T08:36:42.033Z",
          "content": "<p>LSA features were computed on the original data :) </p>",
          "rawMarkdown": "LSA features were computed on the original data :) ",
          "votes": 2
        },
        {
          "id": 326731,
          "postDate": "2018-05-10T08:38:22.307Z",
          "content": "<p>Thanks,.</p>\n\n<p>You almost make me try to rerun my stuff on downsampled data ;)  </p>\n\n<p>Really neat solution, very impressive.  The gap between us is well justified.</p>",
          "rawMarkdown": "Thanks,.\n\nYou almost make me try to rerun my stuff on downsampled data ;)  \n\nReally neat solution, very impressive.  The gap between us is well justified.",
          "votes": 2
        },
        {
          "id": 326961,
          "postDate": "2018-05-10T15:14:50.697Z",
          "content": "<p>Agree. It's a very neat solution, just like a clever and beautiful solution to a mathematics problem.</p>",
          "rawMarkdown": "Agree. It's a very neat solution, just like a clever and beautiful solution to a mathematics problem."
        },
        {
          "id": 327295,
          "postDate": "2018-05-11T07:49:58.520Z",
          "content": "<p>Me too! I feel stupider because I've implemented downsampling of negative 3 weeks ago but forgot to run it until the competition ended. XD</p>",
          "rawMarkdown": "Me too! I feel stupider because I've implemented downsampling of negative 3 weeks ago but forgot to run it until the competition ended. XD"
        }
      ]
    },
    {
      "id": 327266,
      "postDate": "2018-05-11T06:45:08.493Z",
      "content": "<p>Do you use sklearn.decomposition.LatentDirichletAllocation ? or others?  It's very slow. How to train it quickly?</p>",
      "rawMarkdown": "Do you use sklearn.decomposition.LatentDirichletAllocation ? or others?  It's very slow. How to train it quickly?",
      "votes": 1,
      "replies": [
        {
          "id": 328142,
          "postDate": "2018-05-13T13:18:27.490Z",
          "content": "<p>Yes, we used sklearn.decomposition.LatentDirichletAllocation. If my memory is correct, it takes about one hour for calculation on an (ip, app) pair. It was acceptable because we stored each feature in a file once we  computed it and we don't have to run LDA for each training. </p>",
          "rawMarkdown": "Yes, we used sklearn.decomposition.LatentDirichletAllocation. If my memory is correct, it takes about one hour for calculation on an (ip, app) pair. It was acceptable because we stored each feature in a file once we  computed it and we don't have to run LDA for each training. "
        },
        {
          "id": 328154,
          "postDate": "2018-05-13T13:36:13.703Z",
          "content": "<p>@xuan Class LDA has a parameter n_jobs. Have you tried it?\n<a href=\"/flowlight\">@flowlight</a>. Congrats and thanks for your sharing. Choosing a proper number of components often gave me a headache. I was wondering if there exist an empirical relationship between n_components and the shape (rows and columns) of the matrix. Or could you please tell me why you choose 5 for it? Is it just chosen arbitrarily or tuned?</p>",
          "rawMarkdown": "@xuan Class LDA has a parameter n_jobs. Have you tried it?\n@flowlight. Congrats and thanks for your sharing. Choosing a proper number of components often gave me a headache. I was wondering if there exist an empirical relationship between n_components and the shape (rows and columns) of the matrix. Or could you please tell me why you choose 5 for it? Is it just chosen arbitrarily or tuned?"
        },
        {
          "id": 328199,
          "postDate": "2018-05-13T16:26:13.883Z",
          "content": "<p>There are no strong reason why we chose 5 as n_components. Actually, we also added LDA with n_components = 20 before the deadline of this competition. </p>",
          "rawMarkdown": "There are no strong reason why we chose 5 as n_components. Actually, we also added LDA with n_components = 20 before the deadline of this competition. "
        },
        {
          "id": 328319,
          "postDate": "2018-05-14T01:04:14.247Z",
          "content": "<p>Thanks a lot for your reply!</p>",
          "rawMarkdown": "Thanks a lot for your reply!"
        },
        {
          "id": 329285,
          "postDate": "2018-05-16T06:35:09.660Z",
          "content": "<p>n_jobs = -1, can quikly a lot,I try</p>",
          "rawMarkdown": "n_jobs = -1, can quikly a lot,I try"
        }
      ]
    },
    {
      "id": 327033,
      "postDate": "2018-05-10T17:56:24.310Z",
      "content": "<p>Congratulations!Thanks for sharing.Also want to know why \"removed all raw categorical features except app\" work....</p>",
      "rawMarkdown": "Congratulations!Thanks for sharing.Also want to know why \"removed all raw categorical features except app\" work....",
      "votes": 1,
      "replies": [
        {
          "id": 327870,
          "postDate": "2018-05-12T18:55:02.070Z",
          "content": "<p>Removing all but app and os worked for me.  Don't know why either.</p>",
          "rawMarkdown": "Removing all but app and os worked for me.  Don't know why either."
        },
        {
          "id": 327918,
          "postDate": "2018-05-12T22:23:35.913Z",
          "content": "<p>My guess:</p>\n\n<p>The categorical features are simply labels and the numbers represented by the labels are more or less arbitrary. The model is successfully finding signal in the engineered features, such as count by ip, but the original categorical features are just adding more noise to the model, and when training, the model's learning of the actual relationships between engineered features and is_attributed is being diluted by the categorical features. </p>\n\n<p>There might be some information in the categorical features, maybe some of the app ids that have higher assigned numbers are correlated is is_attributed, but that is possibly over fitting?</p>\n\n<p>For reference, my model had very few engineered features and I struggled greatly to find any signal. </p>",
          "rawMarkdown": "My guess:\n\nThe categorical features are simply labels and the numbers represented by the labels are more or less arbitrary. The model is successfully finding signal in the engineered features, such as count by ip, but the original categorical features are just adding more noise to the model, and when training, the model's learning of the actual relationships between engineered features and is_attributed is being diluted by the categorical features. \n\nThere might be some information in the categorical features, maybe some of the app ids that have higher assigned numbers are correlated is is_attributed, but that is possibly over fitting?\n\nFor reference, my model had very few engineered features and I struggled greatly to find any signal. "
        },
        {
          "id": 327977,
          "postDate": "2018-05-13T04:07:21.703Z",
          "content": "<p>When you use a feature as a categorical in LightGBM the ordering of the numbers does not matter.</p>",
          "rawMarkdown": "When you use a feature as a categorical in LightGBM the ordering of the numbers does not matter.",
          "votes": 1
        },
        {
          "id": 328030,
          "postDate": "2018-05-13T07:26:32.317Z",
          "content": "<p>My apologies, I thought I was reading about neural networks</p>",
          "rawMarkdown": "My apologies, I thought I was reading about neural networks"
        },
        {
          "id": 329853,
          "postDate": "2018-05-17T09:43:48.857Z",
          "content": "<p>Maybe raw categorical features information are captured by the more complex features. For example,  ip_dev_os and dev or os may be higly correlated. Can this increase the noise instead of adding meaningful predicting value?</p>",
          "rawMarkdown": "Maybe raw categorical features information are captured by the more complex features. For example,  ip_dev_os and dev or os may be higly correlated. Can this increase the noise instead of adding meaningful predicting value?"
        }
      ]
    },
    {
      "id": 386003,
      "postDate": "2018-09-12T00:29:45.827Z",
      "content": "<p>The feature engineering by LDA is brilliant . Thanks for sharing. </p>",
      "rawMarkdown": "The feature engineering by LDA is brilliant . Thanks for sharing. ",
      "votes": 2
    },
    {
      "id": 368871,
      "postDate": "2018-08-11T05:49:17.357Z",
      "content": "<p>Hi there, just a quick question. What is a bagged LGB model? And did you use different random seeds for k-fold or not when creating seven LGB models? Many thanks!</p>",
      "rawMarkdown": "Hi there, just a quick question. What is a bagged LGB model? And did you use different random seeds for k-fold or not when creating seven LGB models? Many thanks!",
      "votes": 2
    },
    {
      "id": 336106,
      "postDate": "2018-05-31T04:26:19.170Z",
      "content": "<p><a href=\"/flowlight\">@flowlight</a>'s repository:\n<a href=\"https://github.com/flowlight0/talkingdata-adtracking-fraud-detection\">https://github.com/flowlight0/talkingdata-adtracking-fraud-detection</a></p>",
      "rawMarkdown": "@flowlight's repository:\nhttps://github.com/flowlight0/talkingdata-adtracking-fraud-detection\n",
      "votes": 2,
      "replies": [
        {
          "id": 336798,
          "postDate": "2018-06-01T09:53:31.933Z",
          "content": "<p>Thanks for attaching this link. I have forgotten to post a comment about my repo. </p>",
          "rawMarkdown": "Thanks for attaching this link. I have forgotten to post a comment about my repo. ",
          "votes": 2
        },
        {
          "id": 336816,
          "postDate": "2018-06-01T10:39:21.540Z",
          "content": "<p>Thank you for sharing your repository. Your creative solution is valuable not only for current but also for future data scientists, researchers and machine learning engineers.  Providing its repository makes it more informative and helpful.</p>",
          "rawMarkdown": "Thank you for sharing your repository. Your creative solution is valuable not only for current but also for future data scientists, researchers and machine learning engineers.  Providing its repository makes it more informative and helpful."
        }
      ]
    },
    {
      "id": 326966,
      "postDate": "2018-05-10T15:22:00.910Z",
      "content": "<blockquote>\n  <ul>\n  <li>average attributed ratio of past clicks</li>\n  </ul>\n</blockquote>\n\n<p>Doesn't this feature cause over-fit? Do you confirm or are you sure this feature really improve the score?\nCould you provide the feature importance of your trained model?</p>",
      "rawMarkdown": "&gt; - average attributed ratio of past clicks\n\nDoesn't this feature cause over-fit? Do you confirm or are you sure this feature really improve the score?\nCould you provide the feature importance of your trained model?\n",
      "votes": 2,
      "replies": [
        {
          "id": 326972,
          "postDate": "2018-05-10T15:30:11.180Z",
          "content": "<p>Maybe I don't understand the sentence at all. What does \"past clicks\" mean?\nCould you show a pseudo code like following?:</p>\n\n<blockquote>\n  <p>test_and_train_df.groupby(group)['is_attributed'].mean()</p>\n</blockquote>",
          "rawMarkdown": "Maybe I don't understand the sentence at all. What does \"past clicks\" mean?\nCould you show a pseudo code like following?:\n\n&gt;  test_and_train_df.groupby(group)['is_attributed'].mean()\n",
          "votes": 1
        },
        {
          "id": 327100,
          "postDate": "2018-05-10T20:17:58.547Z",
          "content": "<p>I asked in other forum and got an answer. I show the idea as pseudo code.</p>\n\n<pre><code>s = '{}_mean_is_attributed'.format('_'.join(group))\nday9.join(day8.groupby(group)['is_attributed'].mean().rename(s), on=group)\nday10.join(day9.groupby(group)['is_attributed'].mean().rename(s), on=group)\nday11.join(day10.groupby(group)['is_attributed'].mean().rename(s), on=group)\n</code></pre>\n\n<p>Do I understand correctly?</p>",
          "rawMarkdown": "I asked in other forum and got an answer. I show the idea as pseudo code.\n\n    s = '{}_mean_is_attributed'.format('_'.join(group))\n    day9.join(day8.groupby(group)['is_attributed'].mean().rename(s), on=group)\n    day10.join(day9.groupby(group)['is_attributed'].mean().rename(s), on=group)\n    day11.join(day10.groupby(group)['is_attributed'].mean().rename(s), on=group)\n\nDo I understand correctly?\n"
        },
        {
          "id": 327293,
          "postDate": "2018-05-11T07:43:10.760Z",
          "content": "<p>I tried this feature at first but afterwards I found a better way - using all data excluding the current day to get the attributed rates, which outperforms the one when you just use the data of the previous day. And I'm sure this won't cause overfitting, at least not for me.</p>",
          "rawMarkdown": "I tried this feature at first but afterwards I found a better way - using all data excluding the current day to get the attributed rates, which outperforms the one when you just use the data of the previous day. And I'm sure this won't cause overfitting, at least not for me.",
          "votes": 3
        },
        {
          "id": 328150,
          "postDate": "2018-05-13T13:27:07.870Z",
          "content": "<p>For each click, we calculated average attributed ratio by using clicks that re more than one day before that click. I wanted to make distributions of these features same in both of train and test data. </p>",
          "rawMarkdown": "For each click, we calculated average attributed ratio by using clicks that re more than one day before that click. I wanted to make distributions of these features same in both of train and test data. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1674903,
      "postDate": "2022-02-03T21:24:20.890Z",
      "content": "<p>Very good and clear guidelines indeed!)</p>",
      "rawMarkdown": "Very good and clear guidelines indeed!)"
    },
    {
      "id": 436149,
      "postDate": "2018-12-09T17:27:32.023Z",
      "content": "<p>damn, that's something I have been looking for. Perfect for a rookie.\nThank you</p>",
      "rawMarkdown": "damn, that's something I have been looking for. Perfect for a rookie.\nThank you"
    },
    {
      "id": 372452,
      "postDate": "2018-08-19T12:22:25.630Z",
      "content": "<p>The LDA move is ingenious. Hope I can get the idea and use it somewhere else too.</p>",
      "rawMarkdown": "The LDA move is ingenious. Hope I can get the idea and use it somewhere else too."
    },
    {
      "id": 331424,
      "postDate": "2018-05-21T07:42:18.070Z",
      "content": "<p>Insightful solution, thanks for sharing!</p>",
      "rawMarkdown": "Insightful solution, thanks for sharing!"
    },
    {
      "id": 330683,
      "postDate": "2018-05-19T12:34:10.807Z",
      "content": "<p>Congratulations! A noob and stupid question, why the down-sampling makes such a difference? Are you using the raw model trained with the balanced dataset? Or are you performing some adjustment to take into account this down-sampling (weighting examples, adjusting the trained model, etc)?</p>",
      "rawMarkdown": "Congratulations! A noob and stupid question, why the down-sampling makes such a difference? Are you using the raw model trained with the balanced dataset? Or are you performing some adjustment to take into account this down-sampling (weighting examples, adjusting the trained model, etc)?"
    },
    {
      "id": 330515,
      "postDate": "2018-05-19T03:03:53.790Z",
      "content": "<p>Congrats. Great work.</p>",
      "rawMarkdown": "Congrats. Great work."
    },
    {
      "id": 330292,
      "postDate": "2018-05-18T13:53:41.727Z",
      "content": "<p>Thanks for sharing. It's a great learning experience!</p>",
      "rawMarkdown": "Thanks for sharing. It's a great learning experience!"
    },
    {
      "id": 330079,
      "postDate": "2018-05-18T02:10:25.867Z",
      "content": "<p>Thanks for sharing! Learn a lot from negative down-sampling and categorical feature embedding! Cool jobs!</p>",
      "rawMarkdown": "Thanks for sharing! Learn a lot from negative down-sampling and categorical feature embedding! Cool jobs!"
    },
    {
      "id": 329746,
      "postDate": "2018-05-17T05:16:52.597Z",
      "content": "<p>Congrats. Great work. Well deserved. </p>",
      "rawMarkdown": "Congrats. Great work. Well deserved. "
    },
    {
      "id": 329341,
      "postDate": "2018-05-16T09:22:29.083Z",
      "content": "<p>great solution</p>",
      "rawMarkdown": "great solution"
    },
    {
      "id": 329241,
      "postDate": "2018-05-16T04:34:46.700Z",
      "content": "<p>Thanks for sharing. Few questions</p>\n\n<ol>\n<li>To calculate \"click count within next one/six hours\" feature, did you use test_supplement.csv.zip file?</li>\n<li>For LDA/NMF/LSA related features, could you specify which raw feature is treated as word? I would guess in most cases \"ip\" will be treated as document? </li>\n</ol>",
      "rawMarkdown": "Thanks for sharing. Few questions\n\n1. To calculate \"click count within next one/six hours\" feature, did you use test_supplement.csv.zip file?\n2. For LDA/NMF/LSA related features, could you specify which raw feature is treated as word? I would guess in most cases \"ip\" will be treated as document? "
    },
    {
      "id": 328719,
      "postDate": "2018-05-15T00:14:24.530Z",
      "content": "<p>I am glad I saw this post and learn something I haven't even thought of before. Great work. Congrats!!</p>",
      "rawMarkdown": "I am glad I saw this post and learn something I haven't even thought of before. Great work. Congrats!!"
    },
    {
      "id": 328475,
      "postDate": "2018-05-14T12:18:34.257Z",
      "content": "<p>Nice work</p>",
      "rawMarkdown": "Nice work"
    },
    {
      "id": 328394,
      "postDate": "2018-05-14T07:32:35.227Z",
      "content": "<p>Congrats！ and thanks for sharing your brilliant solutions.</p>\n\n<p>I am a little confused about  when you said “we created higher dimensional LDA features and features that try to address the duplicate sample problem”.</p>\n\n<ol>\n<li>by saying \"higher dimensional\", do you mean some matrix factoraization over combined  categorical features or sth else?</li>\n<li>i dont quite understand “duplicate sample problem” and how it would negatively effect the model. </li>\n</ol>\n\n<p>Enjoy the wins and i am looking forwards to your reply!</p>",
      "rawMarkdown": "Congrats！ and thanks for sharing your brilliant solutions.\n\nI am a little confused about  when you said “we created higher dimensional LDA features and features that try to address the duplicate sample problem”.\n\n1. by saying \"higher dimensional\", do you mean some matrix factoraization over combined  categorical features or sth else?\n2. i dont quite understand “duplicate sample problem” and how it would negatively effect the model. \n\nEnjoy the wins and i am looking forwards to your reply!"
    },
    {
      "id": 328340,
      "postDate": "2018-05-14T02:46:21.483Z",
      "content": "<p>Cool stuff lad</p>",
      "rawMarkdown": "Cool stuff lad"
    },
    {
      "id": 328096,
      "postDate": "2018-05-13T10:35:45.397Z",
      "content": "<p>Congrats and Thanks for sharing your strategy!</p>",
      "rawMarkdown": "Congrats and Thanks for sharing your strategy!"
    },
    {
      "id": 327951,
      "postDate": "2018-05-13T02:34:59.637Z",
      "content": "<p>Congtats！Thank you for your wonderful Insights!</p>",
      "rawMarkdown": "Congtats！Thank you for your wonderful Insights!"
    },
    {
      "id": 327947,
      "postDate": "2018-05-13T02:05:19.730Z",
      "content": "<p>one question about datetime features, did you guys use click_time or attributed time? or perhaps both? </p>",
      "rawMarkdown": "one question about datetime features, did you guys use click_time or attributed time? or perhaps both? ",
      "replies": [
        {
          "id": 328145,
          "postDate": "2018-05-13T13:20:48.847Z",
          "content": "<p>We used only click_time. </p>",
          "rawMarkdown": "We used only click_time. "
        }
      ]
    },
    {
      "id": 327709,
      "postDate": "2018-05-12T08:40:17.743Z",
      "content": "<p>Thanks for posting your solution. I am learning so much from the top solutions. Wonderful Insights</p>",
      "rawMarkdown": "Thanks for posting your solution. I am learning so much from the top solutions. Wonderful Insights"
    },
    {
      "id": 327657,
      "postDate": "2018-05-12T05:28:52.587Z",
      "content": "<p>\"we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)\"</p>\n\n<p>just wondering why is (2^5)-1 here? </p>",
      "rawMarkdown": " \"we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)\"\n\njust wondering why is (2^5)-1 here? ",
      "replies": [
        {
          "id": 327674,
          "postDate": "2018-05-12T06:45:05.137Z",
          "content": "<p>binominal theorem</p>\n\n<blockquote>\n<pre><code>\\binom{5}{0} + \\binom{5}{1} + \\binom{5}{2} + \\binom{5}{3} + \\binom{5}{4} + \\binom{5}{5} \n- \\binom{5}{0}. \n=2^5-1\n</code></pre>\n</blockquote>\n\n<p>don't know if I get the point.</p>\n\n<p>Sorry, but could someone please tell me how to write a math formula in plain text format?</p>",
          "rawMarkdown": "binominal theorem\n\n&gt;     \\binom{5}{0} + \\binom{5}{1} + \\binom{5}{2} + \\binom{5}{3} + \\binom{5}{4} + \\binom{5}{5} \n&gt;     - \\binom{5}{0}. \n&gt;     =2^5-1\n\ndon't know if I get the point.\n\nSorry, but could someone please tell me how to write a math formula in plain text format?"
        },
        {
          "id": 327727,
          "postDate": "2018-05-12T09:27:51.397Z",
          "content": "<p>@Kelvin, There are 2^5 subsets of a set of 5 elements.  If you discard the empty set then you get 2^5 - 1.</p>\n\n<p>@Rayarrow I don't know how to use Latex or math markup here.</p>",
          "rawMarkdown": "@Kelvin, There are 2^5 subsets of a set of 5 elements.  If you discard the empty set then you get 2^5 - 1.\n\n@Rayarrow I don't know how to use Latex or math markup here.",
          "votes": 3
        },
        {
          "id": 327730,
          "postDate": "2018-05-12T09:32:34.963Z",
          "content": "<p>Thanks anyway for your reply.</p>",
          "rawMarkdown": "Thanks anyway for your reply.",
          "votes": 1
        },
        {
          "id": 327845,
          "postDate": "2018-05-12T16:28:28.760Z",
          "content": "<p>that's the part im little confused, five raw categorical features are ip, os, app, channel, device, what 5 features of each come from and how come there is one empty set ? maybe I missed some info here?  Sorry, I'm just trying to understand here.</p>",
          "rawMarkdown": "that's the part im little confused, five raw categorical features are ip, os, app, channel, device, what 5 features of each come from and how come there is one empty set ? maybe I missed some info here?  Sorry, I'm just trying to understand here."
        },
        {
          "id": 327853,
          "postDate": "2018-05-12T17:02:15.793Z",
          "content": "<p>@Kelvin, some more details: it's called the <a href=\"https://en.wikipedia.org/wiki/Power_set\">power set</a>, for n elements it's size is 2^n... Here's some code to see it explicitly:</p>\n\n<pre><code>from itertools import chain, combinations\n\ndef powerset(iterable):\n    s = list(iterable)\n    return chain.from_iterable(combinations(s, r) for r in range(len(s)+1))\n\nlist(enumerate(powerset(cols)))\n</code></pre>\n\n<p>They did it <strong><em>the Kaggle way</em></strong> and generated the same feature set for all of 1..31 :)</p>\n\n<pre><code>[(0, ()),\n (1, ('app',)),\n (2, ('channel',)),\n (3, ('ip',)),\n (4, ('device',)),\n (5, ('os',)),\n (6, ('app', 'channel')),\n (7, ('app', 'ip')),\n (8, ('app', 'device')),\n (9, ('app', 'os')),\n (10, ('channel', 'ip')),\n (11, ('channel', 'device')),\n (12, ('channel', 'os')),\n (13, ('ip', 'device')),\n (14, ('ip', 'os')),\n (15, ('device', 'os')),\n (16, ('app', 'channel', 'ip')),\n (17, ('app', 'channel', 'device')),\n (18, ('app', 'channel', 'os')),\n (19, ('app', 'ip', 'device')),\n (20, ('app', 'ip', 'os')),\n (21, ('app', 'device', 'os')),\n (22, ('channel', 'ip', 'device')),\n (23, ('channel', 'ip', 'os')),\n (24, ('channel', 'device', 'os')),\n (25, ('ip', 'device', 'os')),\n (26, ('app', 'channel', 'ip', 'device')),\n (27, ('app', 'channel', 'ip', 'os')),\n (28, ('app', 'channel', 'device', 'os')),\n (29, ('app', 'ip', 'device', 'os')),\n (30, ('channel', 'ip', 'device', 'os')),\n (31, ('app', 'channel', 'ip', 'device', 'os'))]\n</code></pre>\n\n<p>They also have one of the best team names, <code>['flowlight', 'komaki'].shuffle()</code>, both get <a href=\"https://en.wikipedia.org/wiki/Billing_%28filmmaking%29\">top billing</a>, 50% of the time ;)  Congrats on the win guys!</p>",
          "rawMarkdown": "@Kelvin, some more details: it's called the [power set][1], for n elements it's size is 2^n... Here's some code to see it explicitly:\n\n    from itertools import chain, combinations\n    \n    def powerset(iterable):\n        s = list(iterable)\n        return chain.from_iterable(combinations(s, r) for r in range(len(s)+1))\n    \n    list(enumerate(powerset(cols)))\n\nThey did it ***the Kaggle way*** and generated the same feature set for all of 1..31 :)\n\n\n    [(0, ()),\n     (1, ('app',)),\n     (2, ('channel',)),\n     (3, ('ip',)),\n     (4, ('device',)),\n     (5, ('os',)),\n     (6, ('app', 'channel')),\n     (7, ('app', 'ip')),\n     (8, ('app', 'device')),\n     (9, ('app', 'os')),\n     (10, ('channel', 'ip')),\n     (11, ('channel', 'device')),\n     (12, ('channel', 'os')),\n     (13, ('ip', 'device')),\n     (14, ('ip', 'os')),\n     (15, ('device', 'os')),\n     (16, ('app', 'channel', 'ip')),\n     (17, ('app', 'channel', 'device')),\n     (18, ('app', 'channel', 'os')),\n     (19, ('app', 'ip', 'device')),\n     (20, ('app', 'ip', 'os')),\n     (21, ('app', 'device', 'os')),\n     (22, ('channel', 'ip', 'device')),\n     (23, ('channel', 'ip', 'os')),\n     (24, ('channel', 'device', 'os')),\n     (25, ('ip', 'device', 'os')),\n     (26, ('app', 'channel', 'ip', 'device')),\n     (27, ('app', 'channel', 'ip', 'os')),\n     (28, ('app', 'channel', 'device', 'os')),\n     (29, ('app', 'ip', 'device', 'os')),\n     (30, ('channel', 'ip', 'device', 'os')),\n     (31, ('app', 'channel', 'ip', 'device', 'os'))]\n\n\nThey also have one of the best team names, `['flowlight', 'komaki'].shuffle()`, both get [top billing][2], 50% of the time ;)  Congrats on the win guys!\n\n\n  [1]: https://en.wikipedia.org/wiki/Power_set\n  [2]: https://en.wikipedia.org/wiki/Billing_(filmmaking)",
          "votes": 7
        },
        {
          "id": 327866,
          "postDate": "2018-05-12T18:21:36.290Z",
          "content": "<p>@James\noh i see, they actually list all possible elements in combinations,  i got it now. appreciate it your help!</p>",
          "rawMarkdown": "@James\noh i see, they actually list all possible elements in combinations,  i got it now. appreciate it your help!"
        },
        {
          "id": 329342,
          "postDate": "2018-05-16T09:24:01.070Z",
          "content": "<p>nice</p>",
          "rawMarkdown": "nice"
        }
      ]
    },
    {
      "id": 327598,
      "postDate": "2018-05-12T00:32:33.163Z",
      "content": "<p>Congratulations on the win.</p>\n\n<p>The paper 'A View from the Trenches' references using a 'feature importance weight' to scale down the learning from the positive results, so the network output didn't have a bias toward positive samples, did you implement something similar? </p>\n\n<p>Wouldn't a model trained on oversampled positive data predict more positives? </p>",
      "rawMarkdown": "Congratulations on the win.\n\nThe paper 'A View from the Trenches' references using a 'feature importance weight' to scale down the learning from the positive results, so the network output didn't have a bias toward positive samples, did you implement something similar? \n\nWouldn't a model trained on oversampled positive data predict more positives? ",
      "replies": [
        {
          "id": 328313,
          "postDate": "2018-05-14T00:15:42.533Z",
          "content": "<p>no. Predicting more positives wasn't a problem for AUC, luckily.</p>",
          "rawMarkdown": "no. Predicting more positives wasn't a problem for AUC, luckily."
        }
      ]
    },
    {
      "id": 327205,
      "postDate": "2018-05-11T02:44:21.247Z",
      "content": "<p>Maybe the reason that why removing all raw categorical features works is that your embedding features have captured more information than raw features. Thus raw feature become noise.</p>",
      "rawMarkdown": "Maybe the reason that why removing all raw categorical features works is that your embedding features have captured more information than raw features. Thus raw feature become noise."
    },
    {
      "id": 327199,
      "postDate": "2018-05-11T02:15:26.623Z",
      "content": "<p>Congratulations.Learn more from your solution. Especially categorical features processing method </p>",
      "rawMarkdown": "Congratulations.Learn more from your solution. Especially categorical features processing method "
    },
    {
      "id": 327070,
      "postDate": "2018-05-10T19:21:33.617Z",
      "content": "<p>Congrats! The methods are brilliant! Your methods and features are quite different from us, which is really amazing. I am sure I can learn a lot from you. Thanks!</p>",
      "rawMarkdown": "Congrats! The methods are brilliant! Your methods and features are quite different from us, which is really amazing. I am sure I can learn a lot from you. Thanks!"
    },
    {
      "id": 327036,
      "postDate": "2018-05-10T18:01:38.417Z",
      "content": "<p>Beautiful solution! Thank you very much for sharing and congrats! </p>",
      "rawMarkdown": "Beautiful solution! Thank you very much for sharing and congrats! "
    },
    {
      "id": 326935,
      "postDate": "2018-05-10T14:31:18.643Z",
      "content": "<p>hello，I'm a beginner in kaggle competitions. can you explain how to  created a bunch of features in a brute-force way?\nThanks</p>",
      "rawMarkdown": "hello，I'm a beginner in kaggle competitions. can you explain how to  created a bunch of features in a brute-force way?\nThanks"
    },
    {
      "id": 326815,
      "postDate": "2018-05-10T11:31:29.200Z",
      "content": "<p>Thanks for sharing guys. I admire the efficiency of the downsampling for training, yet the breadth of feature exploration. Nice work!</p>",
      "rawMarkdown": "Thanks for sharing guys. I admire the efficiency of the downsampling for training, yet the breadth of feature exploration. Nice work!"
    },
    {
      "id": 326743,
      "postDate": "2018-05-10T09:00:47.137Z",
      "content": "<p>Thanks and congrats! <br>\nAnd just a quick question: when you use subsample to decrease the size of negative labels, do you select them randomly every time you train model or just use the same sets every time? </p>",
      "rawMarkdown": "Thanks and congrats!  \nAnd just a quick question: when you use subsample to decrease the size of negative labels, do you select them randomly every time you train model or just use the same sets every time? ",
      "replies": [
        {
          "id": 326781,
          "postDate": "2018-05-10T10:26:39.197Z",
          "content": "<p>Thank you. We did the former. We changed our training dataset for each training to get more diversity.</p>",
          "rawMarkdown": "Thank you. We did the former. We changed our training dataset for each training to get more diversity.",
          "votes": 4
        }
      ]
    },
    {
      "id": 326732,
      "postDate": "2018-05-10T08:41:17.423Z",
      "content": "<p>I'm a beginner in kaggle competitions.I‘m wondered how to bag LGB model.Does it use the output of last layer feature transform as the input of current layer?   </p>",
      "rawMarkdown": "I'm a beginner in kaggle competitions.I‘m wondered how to bag LGB model.Does it use the output of last layer feature transform as the input of current layer?   ",
      "replies": [
        {
          "id": 326741,
          "postDate": "2018-05-10T08:59:03.933Z",
          "content": "<p>Basically we did the following two steps. <br>\n1. download past submissions from Kaggle <br>\n2. Run the following code <br></p>\n\n<pre>submissions = [pd.read_csv(path) for path in downloaded_submission_paths]\nsubmissions_as_rank = [submission.rank() for submission in submissions]\nensemble = np.mean(submissions_as_rank, axis=0)\n</pre>",
          "rawMarkdown": "Basically we did the following two steps. <br>\n1. download past submissions from Kaggle <br>\n2. Run the following code <br>\n<pre>submissions = [pd.read_csv(path) for path in downloaded_submission_paths]\nsubmissions_as_rank = [submission.rank() for submission in submissions]\nensemble = np.mean(submissions_as_rank, axis=0)\n</pre>\n\n",
          "votes": 6
        },
        {
          "id": 327183,
          "postDate": "2018-05-11T01:23:19.027Z",
          "content": "<p>Does ur LGB model use same features when u training? </p>",
          "rawMarkdown": "Does ur LGB model use same features when u training? "
        }
      ]
    },
    {
      "id": 326715,
      "postDate": "2018-05-10T08:12:49.670Z",
      "content": "<p>congtats！awesome work！</p>",
      "rawMarkdown": "congtats！awesome work！"
    },
    {
      "id": 326710,
      "postDate": "2018-05-10T07:54:36.590Z",
      "content": "<p>Congratulations!  I have a question:your features are calculate on the full data or on the down-sample features? i mean your models diversity is not only train data difference but also feature differenct or just train data difference?</p>",
      "rawMarkdown": "Congratulations!  I have a question:your features are calculate on the full data or on the down-sample features? i mean your models diversity is not only train data difference but also feature differenct or just train data difference?",
      "replies": [
        {
          "id": 326727,
          "postDate": "2018-05-10T08:35:48.410Z",
          "content": "<p>Thank you ChinaBoy. All the features were computed on the full data. Negative down-sampling was done only on model training. </p>",
          "rawMarkdown": "Thank you ChinaBoy. All the features were computed on the full data. Negative down-sampling was done only on model training. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 326708,
      "postDate": "2018-05-10T07:52:08.940Z",
      "content": "<p>Awesome! Your score with downsampling made me happy :)</p>",
      "rawMarkdown": "Awesome! Your score with downsampling made me happy :)"
    },
    {
      "id": 326697,
      "postDate": "2018-05-10T07:33:35.203Z",
      "content": "<p>Congrats and thx for sharing, these two papers are really nice.  A question : brute-force way will bring many features, so how do you choose useful feature?</p>",
      "rawMarkdown": "Congrats and thx for sharing, these two papers are really nice.  A question : brute-force way will bring many features, so how do you choose useful feature?",
      "replies": [
        {
          "id": 326724,
          "postDate": "2018-05-10T08:33:31.140Z",
          "content": "<p>Thank you, zr. Both of the two papers are good, but actually the second link was not what I wanted to put. I replaced it with the correct one. Sorry about this. </p>\n\n<p>Thanks to negative down-sampling, the many-features was not a so big problem to us, so we didn't do fine-grained feature selection. For example, when I came up with the LDA features, I added 100 new features to my model. Luckily, it worked very well, so I added all of the 100 features. These 100 features were used for the final submission model without any further testing. </p>\n\n<p>I haven't cared about which one of them contribute the most though ideally I might have had to remove useless features. </p>",
          "rawMarkdown": "Thank you, zr. Both of the two papers are good, but actually the second link was not what I wanted to put. I replaced it with the correct one. Sorry about this. \n\nThanks to negative down-sampling, the many-features was not a so big problem to us, so we didn't do fine-grained feature selection. For example, when I came up with the LDA features, I added 100 new features to my model. Luckily, it worked very well, so I added all of the 100 features. These 100 features were used for the final submission model without any further testing. \n\nI haven't cared about which one of them contribute the most though ideally I might have had to remove useless features. \n",
          "votes": 2
        }
      ]
    },
    {
      "id": 326694,
      "postDate": "2018-05-10T07:26:32.350Z",
      "content": "<p>Congratulations! I really like the idea with class balancing - I tried it in other way round  boosting positive examples.  </p>",
      "rawMarkdown": "Congratulations! I really like the idea with class balancing - I tried it in other way round  boosting positive examples.  "
    },
    {
      "id": 326685,
      "postDate": "2018-05-10T07:17:17.843Z",
      "content": "<p>Congrats on the result!  Thanks for sharing.  I am not surprised that relying heavily on matrix factorizations helped a lot.</p>",
      "rawMarkdown": "Congrats on the result!  Thanks for sharing.  I am not surprised that relying heavily on matrix factorizations helped a lot.",
      "replies": [
        {
          "id": 326704,
          "postDate": "2018-05-10T07:46:54.943Z",
          "content": "<p>Hi @CPMP ,can you explain about the  intuitions of matrix factorizations??I think it is very amazing</p>",
          "rawMarkdown": "Hi @CPMP ,can you explain about the  intuitions of matrix factorizations??I think it is very amazing"
        },
        {
          "id": 326709,
          "postDate": "2018-05-10T07:52:55.247Z",
          "content": "<p>@ChinaBoy, I thought of it when reading a solution from a previous competition, outbrain click prediction probably.  The authors (was it 3 idiots?) said that using the set of apps for a given user helped a lot.  In our competition here, representing this set would be too large, I therefore tried to find a way to get an approximation.  That's how I used tSVD on the userxapp matrix.  Note that tSVD is exactly the same model as the lsa model use by Komaki's team.  The only difference is how data is input to it.</p>",
          "rawMarkdown": "@ChinaBoy, I thought of it when reading a solution from a previous competition, outbrain click prediction probably.  The authors (was it 3 idiots?) said that using the set of apps for a given user helped a lot.  In our competition here, representing this set would be too large, I therefore tried to find a way to get an approximation.  That's how I used tSVD on the userxapp matrix.  Note that tSVD is exactly the same model as the lsa model use by Komaki's team.  The only difference is how data is input to it.",
          "votes": 5
        },
        {
          "id": 326711,
          "postDate": "2018-05-10T07:56:13.187Z",
          "content": "<p>yes.i also think so~thank you for your reply</p>",
          "rawMarkdown": "yes.i also think so~thank you for your reply"
        }
      ]
    },
    {
      "id": 627607,
      "postDate": "2019-09-16T07:07:58.857Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 330050,
      "postDate": "2018-05-17T23:00:49Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1071619,
      "postDate": "2020-11-07T05:52:35.600Z",
      "content": "<p>Thank you for sharing your approach 😃</p>",
      "rawMarkdown": "Thank you for sharing your approach 😃"
    },
    {
      "id": 436248,
      "postDate": "2018-12-10T01:25:41.773Z",
      "content": "<p>thank you</p>",
      "rawMarkdown": "thank you"
    },
    {
      "id": 401111,
      "postDate": "2018-10-09T12:57:56.033Z",
      "content": "<p>Thank you for sharing. I will try it.</p>",
      "rawMarkdown": "Thank you for sharing. I will try it."
    },
    {
      "id": 332137,
      "postDate": "2018-05-22T15:54:22.030Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 330726,
      "postDate": "2018-05-19T14:53:33.800Z",
      "content": "<p>Congratulations!Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations!Thanks for sharing."
    },
    {
      "id": 328344,
      "postDate": "2018-05-14T03:00:09.620Z",
      "content": "<p>Thank you for sharing! </p>",
      "rawMarkdown": "Thank you for sharing! "
    }
  ],
  "comments": [
    {
      "id": 326747,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "2018-05-10T09:09:32.277000",
      "content": "<p>Congrats,a very good solution,your 1st place is well deserved,thanks for sharing. I think if we can combine the top teams' solutions ,we will get a very high score.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 326757,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T09:23:19.427000",
          "content": "<p>I am not sure I qualify as one of the top teams, but I shared my 0.98379 private solution <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423\">here</a>.  Maybe others could do the same so that we can blend?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 326761,
          "author_name": "bestfitting",
          "author_url": "",
          "post_date": "2018-05-10T09:30:09.453000",
          "content": "<p>Absolutely should include you  :)<br>To get top 50 is not so easy in this competition.<br></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 327078,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-10T19:31:14.327000",
          "content": "<p>I feel the same. I am very impressed to see so many wonderful ideas that I have never thought of. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 326793,
      "author_name": "Pranav Pandya",
      "author_url": "",
      "post_date": "2018-05-10T10:55:21.323000",
      "content": "<p>Congratulations <a href=\"/ckomaki\">@ckomaki</a> and <a href=\"/flowlight\">@flowlight</a>. Excellent approach and very well deserved win. And a big thank you for the mention :) This actually made my day! </p>",
      "votes": 8,
      "replies": [
        {
          "id": 326932,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T14:27:23.777000",
          "content": "<p>@Pranav, It is a pity to see how Dirk share impacted you.  Most people reused your shared kernels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326946,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-05-10T14:54:21.077000",
          "content": "<p>That's alrite @CPMP. Though Dirk impacted me too but I also blame myself for making stupid selections on judgement day :) But honestly, this competition was really a great learning experience for me specifically in feature engineering part. </p>\n\n<p>Shared learning from community actually helped me a lot and I'm sure that this will contribute positively for me in next competitions :) </p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 326714,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2018-05-10T08:07:04.547000",
      "content": "<p>fantastic solution, thank you very much for sharing - Matrix factorization for feature interaction will be a main take away for me from this competition </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 326705,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-10T07:49:28.920000",
      "content": "<p>I must say that I feel stupid for not having tested downsampling of negative.  I usually try it on very imbalanced datasets.  Well done!</p>\n\n<p>Are you computing your lsa models on downsampled data or on the original one?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 326729,
          "author_name": "Komaki",
          "author_url": "",
          "post_date": "2018-05-10T08:36:42.033000",
          "content": "<p>LSA features were computed on the original data :) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326731,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-10T08:38:22.307000",
          "content": "<p>Thanks,.</p>\n\n<p>You almost make me try to rerun my stuff on downsampled data ;)  </p>\n\n<p>Really neat solution, very impressive.  The gap between us is well justified.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326961,
          "author_name": "seaguII",
          "author_url": "",
          "post_date": "2018-05-10T15:14:50.697000",
          "content": "<p>Agree. It's a very neat solution, just like a clever and beautiful solution to a mathematics problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327295,
          "author_name": "Rayarrow",
          "author_url": "",
          "post_date": "2018-05-11T07:49:58.520000",
          "content": "<p>Me too! I feel stupider because I've implemented downsampling of negative 3 weeks ago but forgot to run it until the competition ended. XD</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 327266,
      "author_name": "xuan",
      "author_url": "",
      "post_date": "2018-05-11T06:45:08.493000",
      "content": "<p>Do you use sklearn.decomposition.LatentDirichletAllocation ? or others?  It's very slow. How to train it quickly?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 328142,
          "author_name": "flowlight",
          "author_url": "",
          "post_date": "2018-05-13T13:18:27.490000",
          "content": "<p>Yes, we used sklearn.decomposition.LatentDirichletAllocation. If my memory is correct, it takes about one hour for calculation on an (ip, app) pair. It was acceptable because we stored each feature in a file once we  computed it and we don't have to run LDA for each training. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 328154,
          "author_name": "Rayarrow",
          "author_url": "",
          "post_date": "2018-05-13T13:36:13.703000",
          "content": "<p>@xuan Class LDA has a parameter n_jobs. Have you tried it?\n<a href=\"/flowlight\">@flowlight</a>. Congrats and thanks for your sharing. Choosing a proper number of components often gave me a headache. I was wondering if there exist an empirical relationship between n_components and the shape (rows and columns) of the matrix. Or could you please tell me why you choose 5 for it? Is it just chosen arbitrarily or tuned?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 328199,
          "author_name": "flowlight",
          "author_url": "",
          "post_date": "2018-05-13T16:26:13.883000",
          "content": "<p>There are no strong reason why we chose 5 as n_components. Actually, we also added LDA with n_components = 20 before the deadline of this competition. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 328319,
          "author_name": "Rayarrow",
          "author_url": "",
          "post_date": "2018-05-14T01:04:14.247000",
          "content": "<p>Thanks a lot for your reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329285,
          "author_name": "xuan",
          "author_url": "",
          "post_date": "2018-05-16T06:35:09.660000",
          "content": "<p>n_jobs = -1, can quikly a lot,I try</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 327033,
      "author_name": "plantsgo",
      "author_url": "",
      "post_date": "2018-05-10T17:56:24.310000",
      "content": "<p>Congratulations!Thanks for sharing.Also want to know why \"removed all raw categorical features except app\" work....</p>",
      "votes": 1,
      "replies": [
        {
          "id": 327870,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-12T18:55:02.070000",
          "content": "<p>Removing all but app and os worked for me.  Don't know why either.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327918,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-05-12T22:23:35.913000",
          "content": "<p>My guess:</p>\n\n<p>The categorical features are simply labels and the numbers represented by the labels are more or less arbitrary. The model is successfully finding signal in the engineered features, such as count by ip, but the original categorical features are just adding more noise to the model, and when training, the model's learning of the actual relationships between engineered features and is_attributed is being diluted by the categorical features. </p>\n\n<p>There might be some information in the categorical features, maybe some of the app ids that have higher assigned numbers are correlated is is_attributed, but that is possibly over fitting?</p>\n\n<p>For reference, my model had very few engineered features and I struggled greatly to find any signal. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327977,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-13T04:07:21.703000",
          "content": "<p>When you use a feature as a categorical in LightGBM the ordering of the numbers does not matter.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 328030,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-05-13T07:26:32.317000",
          "content": "<p>My apologies, I thought I was reading about neural networks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329853,
          "author_name": "Kate Gallo",
          "author_url": "",
          "post_date": "2018-05-17T09:43:48.857000",
          "content": "<p>Maybe raw categorical features information are captured by the more complex features. For example,  ip_dev_os and dev or os may be higly correlated. Can this increase the noise instead of adding meaningful predicting value?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 386003,
      "author_name": "FredGu",
      "author_url": "",
      "post_date": "2018-09-12T00:29:45.827000",
      "content": "<p>The feature engineering by LDA is brilliant . Thanks for sharing. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 368871,
      "author_name": "seaguII",
      "author_url": "",
      "post_date": "2018-08-11T05:49:17.357000",
      "content": "<p>Hi there, just a quick question. What is a bagged LGB model? And did you use different random seeds for k-fold or not when creating seven LGB models? Many thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 336106,
      "author_name": "Fujii Hironori",
      "author_url": "",
      "post_date": "2018-05-31T04:26:19.170000",
      "content": "<p><a href=\"/flowlight\">@flowlight</a>'s repository:\n<a href=\"https://github.com/flowlight0/talkingdata-adtracking-fraud-detection\">https://github.com/flowlight0/talkingdata-adtracking-fraud-detection</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 336798,
          "author_name": "flowlight",
          "author_url": "",
          "post_date": "2018-06-01T09:53:31.933000",
          "content": "<p>Thanks for attaching this link. I have forgotten to post a comment about my repo. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 336816,
          "author_name": "Fujii Hironori",
          "author_url": "",
          "post_date": "2018-06-01T10:39:21.540000",
          "content": "<p>Thank you for sharing your repository. Your creative solution is valuable not only for current but also for future data scientists, researchers and machine learning engineers.  Providing its repository makes it more informative and helpful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326966,
      "author_name": "Fujii Hironori",
      "author_url": "",
      "post_date": "2018-05-10T15:22:00.910000",
      "content": "<blockquote>\n  <ul>\n  <li>average attributed ratio of past clicks</li>\n  </ul>\n</blockquote>\n\n<p>Doesn't this feature cause over-fit? Do you confirm or are you sure this feature really improve the score?\nCould you provide the feature importance of your trained model?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 326972,
          "author_name": "Fujii Hironori",
          "author_url": "",
          "post_date": "2018-05-10T15:30:11.180000",
          "content": "<p>Maybe I don't understand the sentence at all. What does \"past clicks\" mean?\nCould you show a pseudo code like following?:</p>\n\n<blockquote>\n  <p>test_and_train_df.groupby(group)['is_attributed'].mean()</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 327100,
          "author_name": "Fujii Hironori",
          "author_url": "",
          "post_date": "2018-05-10T20:17:58.547000",
          "content": "<p>I asked in other forum and got an answer. I show the idea as pseudo code.</p>\n\n<pre><code>s = '{}_mean_is_attributed'.format('_'.join(group))\nday9.join(day8.groupby(group)['is_attributed'].mean().rename(s), on=group)\nday10.join(day9.groupby(group)['is_attributed'].mean().rename(s), on=group)\nday11.join(day10.groupby(group)['is_attributed'].mean().rename(s), on=group)\n</code></pre>\n\n<p>Do I understand correctly?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327293,
          "author_name": "Rayarrow",
          "author_url": "",
          "post_date": "2018-05-11T07:43:10.760000",
          "content": "<p>I tried this feature at first but afterwards I found a better way - using all data excluding the current day to get the attributed rates, which outperforms the one when you just use the data of the previous day. And I'm sure this won't cause overfitting, at least not for me.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 328150,
          "author_name": "flowlight",
          "author_url": "",
          "post_date": "2018-05-13T13:27:07.870000",
          "content": "<p>For each click, we calculated average attributed ratio by using clicks that re more than one day before that click. I wanted to make distributions of these features same in both of train and test data. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1674903,
      "author_name": "Kirill Semenov",
      "author_url": "",
      "post_date": "2022-02-03T21:24:20.890000",
      "content": "<p>Very good and clear guidelines indeed!)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436149,
      "author_name": "Ievgen Ishchuk",
      "author_url": "",
      "post_date": "2018-12-09T17:27:32.023000",
      "content": "<p>damn, that's something I have been looking for. Perfect for a rookie.\nThank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 372452,
      "author_name": "dgg32",
      "author_url": "",
      "post_date": "2018-08-19T12:22:25.630000",
      "content": "<p>The LDA move is ingenious. Hope I can get the idea and use it somewhere else too.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 331424,
      "author_name": "datasmith",
      "author_url": "",
      "post_date": "2018-05-21T07:42:18.070000",
      "content": "<p>Insightful solution, thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330683,
      "author_name": "Jordi Villar",
      "author_url": "",
      "post_date": "2018-05-19T12:34:10.807000",
      "content": "<p>Congratulations! A noob and stupid question, why the down-sampling makes such a difference? Are you using the raw model trained with the balanced dataset? Or are you performing some adjustment to take into account this down-sampling (weighting examples, adjusting the trained model, etc)?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330515,
      "author_name": "dean1977",
      "author_url": "",
      "post_date": "2018-05-19T03:03:53.790000",
      "content": "<p>Congrats. Great work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330292,
      "author_name": "TaoL",
      "author_url": "",
      "post_date": "2018-05-18T13:53:41.727000",
      "content": "<p>Thanks for sharing. It's a great learning experience!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330079,
      "author_name": "Shawn Xiao",
      "author_url": "",
      "post_date": "2018-05-18T02:10:25.867000",
      "content": "<p>Thanks for sharing! Learn a lot from negative down-sampling and categorical feature embedding! Cool jobs!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329746,
      "author_name": "AmanSrivastava",
      "author_url": "",
      "post_date": "2018-05-17T05:16:52.597000",
      "content": "<p>Congrats. Great work. Well deserved. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329341,
      "author_name": "joejiong",
      "author_url": "",
      "post_date": "2018-05-16T09:22:29.083000",
      "content": "<p>great solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329241,
      "author_name": "Hang",
      "author_url": "",
      "post_date": "2018-05-16T04:34:46.700000",
      "content": "<p>Thanks for sharing. Few questions</p>\n\n<ol>\n<li>To calculate \"click count within next one/six hours\" feature, did you use test_supplement.csv.zip file?</li>\n<li>For LDA/NMF/LSA related features, could you specify which raw feature is treated as word? I would guess in most cases \"ip\" will be treated as document? </li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328719,
      "author_name": "Kevin Liao",
      "author_url": "",
      "post_date": "2018-05-15T00:14:24.530000",
      "content": "<p>I am glad I saw this post and learn something I haven't even thought of before. Great work. Congrats!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328475,
      "author_name": "Ajit Singh",
      "author_url": "",
      "post_date": "2018-05-14T12:18:34.257000",
      "content": "<p>Nice work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328394,
      "author_name": "tiopon",
      "author_url": "",
      "post_date": "2018-05-14T07:32:35.227000",
      "content": "<p>Congrats！ and thanks for sharing your brilliant solutions.</p>\n\n<p>I am a little confused about  when you said “we created higher dimensional LDA features and features that try to address the duplicate sample problem”.</p>\n\n<ol>\n<li>by saying \"higher dimensional\", do you mean some matrix factoraization over combined  categorical features or sth else?</li>\n<li>i dont quite understand “duplicate sample problem” and how it would negatively effect the model. </li>\n</ol>\n\n<p>Enjoy the wins and i am looking forwards to your reply!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328340,
      "author_name": "Tom Skrovan",
      "author_url": "",
      "post_date": "2018-05-14T02:46:21.483000",
      "content": "<p>Cool stuff lad</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328096,
      "author_name": "Kengo Suzuki",
      "author_url": "",
      "post_date": "2018-05-13T10:35:45.397000",
      "content": "<p>Congrats and Thanks for sharing your strategy!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327951,
      "author_name": "s-kage",
      "author_url": "",
      "post_date": "2018-05-13T02:34:59.637000",
      "content": "<p>Congtats！Thank you for your wonderful Insights!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327947,
      "author_name": "Kelvin",
      "author_url": "",
      "post_date": "2018-05-13T02:05:19.730000",
      "content": "<p>one question about datetime features, did you guys use click_time or attributed time? or perhaps both? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 328145,
          "author_name": "flowlight",
          "author_url": "",
          "post_date": "2018-05-13T13:20:48.847000",
          "content": "<p>We used only click_time. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 327709,
      "author_name": "Akif Rehman",
      "author_url": "",
      "post_date": "2018-05-12T08:40:17.743000",
      "content": "<p>Thanks for posting your solution. I am learning so much from the top solutions. Wonderful Insights</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327657,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-12T05:28:52.587000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 327674,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-12T06:45:05.137000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327727,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-12T09:27:51.397000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 327730,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-12T09:32:34.963000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 327845,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-12T16:28:28.760000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 327853,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-12T17:02:15.793000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 327866,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-12T18:21:36.290000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 329342,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-16T09:24:01.070000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 327598,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-12T00:32:33.163000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 328313,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-14T00:15:42.533000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 327205,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-11T02:44:21.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327199,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-11T02:15:26.623000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327070,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T19:21:33.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 327036,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T18:01:38.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326935,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T14:31:18.643000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326815,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T11:31:29.200000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326743,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T09:00:47.137000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 326781,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T10:26:39.197000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 326732,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T08:41:17.423000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 326741,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T08:59:03.933000",
          "content": "",
          "votes": 6,
          "replies": []
        },
        {
          "id": 327183,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-11T01:23:19.027000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326715,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T08:12:49.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326710,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T07:54:36.590000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 326727,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T08:35:48.410000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 326708,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T07:52:08.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326697,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T07:33:35.203000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 326724,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T08:33:31.140000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 326694,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T07:26:32.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326685,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-10T07:17:17.843000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 326704,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T07:46:54.943000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326709,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T07:52:55.247000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 326711,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-10T07:56:13.187000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 627607,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-16T07:07:58.857000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330050,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-17T23:00:49",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1071619,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-07T05:52:35.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436248,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-10T01:25:41.773000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 401111,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-09T12:57:56.033000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 332137,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-22T15:54:22.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 330726,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-19T14:53:33.800000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328344,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-14T03:00:09.620000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "326680": "We'd like to thank TalkingData and Kaggle for organizing this exciting competition. This competition gave us a fantastic opportunity to learn how to deal with very large table data. <br>\n<br>\nHere is our solution. <br>\n<br>\n**Strategy:** <br>\nOur solution heavily depends on negative down-sampling [1, 2], which means we use all positive examples (i.e., is_attributed == 1) and down-sampled negative examples on model training. We down-sampled negative examples such that their size becomes equal to the number of positive ones. It discards about 99.8% of negative examples, but we didn't see much performance deterioration when we tested with our initial features. Moreover, we could get better performance when creating a submission by bagging five predictors trained on five sampled datasets created from different random seeds. This technique allowed us to use hundreds of features while keeping LGB training time less than 30 minutes. <br>\n<br>\n**Features:** <br>\nFirst, we started from features from Kernels. On Thanks pranav84 and other Kagglers for sharing their awesome insights! We did feature engineerings using all the data examples instead of the down-sampled ones. <br>\n- five raw categorical features (ip, os, app, channel, device) <br>\n- time categorical features (day, hour) <br>\n- some count features <br>\nThen, we created a bunch of features in a brute-force way. For each combination of five raw categorical features (ip, os, app, channel, and device), we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features): <br>\n- click count within next one/six hours <br>\n- forward/backward click time delta <br>\n- average attributed ratio of past clicks <br>\nWe didn't do feature selection. We just added all of them to our model. At that point, our LGB model's score was 0.9808. <br>\n<br>\nNext, we tried categorical feature embedding by using LDA/NMF/LSA. Here is the pseudo code to compute LDA topics of IPs related to app. (LDA is latent Dirichlet allocation)<br>\n<pre>apps_of_ip = {}\nfor sample in data_samples:\n  apps_of_ip.setdefault(sample['ip'], []).append(str(sample['app']))\nips = list(apps_of_ip.keys())\napps_as_sentence = [' '.join(apps_of_ip[ip]) for ip in ips]\napps_as_matrix = CountTokenizer().fit_transform(apps_as_sentence)\ntopics_of_ips = LDA(n_components=5).fit_transform(apps_as_matrix)\n</pre>\nWe computed this feature for all the 20 (=5*(5-1)) combinations of 5 raw features and set the topic size to 5. This ended up with 100 new features. We also computed similar features using NMF and PCA, in total 300 new features. 0.9821 with a single LGB. <br>\n<br>\nAfter that, we removed all raw categorical features except app since we supposed embedding features cover information available from them. Surprisingly, this minor change made our public LB score jump up from 0.9821 to 0.9828. Actually, we don't know what causes this significant score improvement. <br>\nBesides features mentioned here, we created higher dimensional LDA features and features that try to address the duplicate sample problem. These features somewhat improve our public LB score. <br>\n<br>\n**Models:** <br>\nWe used day 7 &amp; 8 for training and day 9 for validation, and chose the best number of iterations of LGB. Then, we trained a model on day 7 &amp; 8 &amp; 9 with the obtained number of iterations for creating submission. After we finished feature engineering, flowlight's five-bagged LGB model reached 0.98333 on public LB (and 0.98420 on private LB), which was trained on 646 features. As far as we remember, the memory usage for training this model was less than 100GB (&lt;64GB will be possible with minor code modification). <br>\nI implemented a simple three layer NN model as some kernels do. It scored worse than LGB models by 0.0013 points with 0.005 down-sampling rate at first. Then, I realized it should be trained with more negative data samples, probably. However, we didn't afford to use so many examples because of the massive amount of our features (and my bad implementation, disk space, no GPU, close deadline, etc...). The final three-bagged NN model scored 0.98258 on public LB. <br>\nWe made our final submission with a rank-based weighted averaging. It is composed of seven bagged LGB models and a single bagged NN  It scored 0.98343 on public LB. <br>\n<br>\n<br>\n[1] <a href=\"http://quinonero.net/Publications/predicting-clicks-facebook.pdf\">Practical Lessons from Predicting Clicks on Ads at\nFacebook</a> <br>\n[2] <a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41159.pdf\">Ad Click Prediction: a View from the Trenches</a> <br>\n<br>\nEDIT: Added some extra explanations based on frequently asked questions in comments.",
    "326747": "Congrats,a very good solution,your 1st place is well deserved,thanks for sharing. I think if we can combine the top teams' solutions ,we will get a very high score.",
    "326793": "Congratulations @ckomaki and @flowlight. Excellent approach and very well deserved win. And a big thank you for the mention :) This actually made my day! ",
    "326714": "fantastic solution, thank you very much for sharing - Matrix factorization for feature interaction will be a main take away for me from this competition ",
    "326705": "I must say that I feel stupid for not having tested downsampling of negative.  I usually try it on very imbalanced datasets.  Well done!\n\nAre you computing your lsa models on downsampled data or on the original one?",
    "327266": "Do you use sklearn.decomposition.LatentDirichletAllocation ? or others?  It's very slow. How to train it quickly?",
    "327033": "Congratulations!Thanks for sharing.Also want to know why \"removed all raw categorical features except app\" work....",
    "386003": "The feature engineering by LDA is brilliant . Thanks for sharing. ",
    "368871": "Hi there, just a quick question. What is a bagged LGB model? And did you use different random seeds for k-fold or not when creating seven LGB models? Many thanks!",
    "336106": "@flowlight's repository:\nhttps://github.com/flowlight0/talkingdata-adtracking-fraud-detection\n",
    "326966": "&gt; - average attributed ratio of past clicks\n\nDoesn't this feature cause over-fit? Do you confirm or are you sure this feature really improve the score?\nCould you provide the feature importance of your trained model?\n",
    "1674903": "Very good and clear guidelines indeed!)",
    "436149": "damn, that's something I have been looking for. Perfect for a rookie.\nThank you",
    "372452": "The LDA move is ingenious. Hope I can get the idea and use it somewhere else too.",
    "331424": "Insightful solution, thanks for sharing!",
    "330683": "Congratulations! A noob and stupid question, why the down-sampling makes such a difference? Are you using the raw model trained with the balanced dataset? Or are you performing some adjustment to take into account this down-sampling (weighting examples, adjusting the trained model, etc)?",
    "330515": "Congrats. Great work.",
    "330292": "Thanks for sharing. It's a great learning experience!",
    "330079": "Thanks for sharing! Learn a lot from negative down-sampling and categorical feature embedding! Cool jobs!",
    "329746": "Congrats. Great work. Well deserved. ",
    "329341": "great solution",
    "329241": "Thanks for sharing. Few questions\n\n1. To calculate \"click count within next one/six hours\" feature, did you use test_supplement.csv.zip file?\n2. For LDA/NMF/LSA related features, could you specify which raw feature is treated as word? I would guess in most cases \"ip\" will be treated as document? ",
    "328719": "I am glad I saw this post and learn something I haven't even thought of before. Great work. Congrats!!",
    "328475": "Nice work",
    "328394": "Congrats！ and thanks for sharing your brilliant solutions.\n\nI am a little confused about  when you said “we created higher dimensional LDA features and features that try to address the duplicate sample problem”.\n\n1. by saying \"higher dimensional\", do you mean some matrix factoraization over combined  categorical features or sth else?\n2. i dont quite understand “duplicate sample problem” and how it would negatively effect the model. \n\nEnjoy the wins and i am looking forwards to your reply!",
    "328340": "Cool stuff lad",
    "328096": "Congrats and Thanks for sharing your strategy!",
    "327951": "Congtats！Thank you for your wonderful Insights!",
    "327947": "one question about datetime features, did you guys use click_time or attributed time? or perhaps both? ",
    "327709": "Thanks for posting your solution. I am learning so much from the top solutions. Wonderful Insights",
    "327657": " \"we created the following click series-based feature sets (i.e., each feature set consists of 31 (=(2^5) - 1) features)\"\n\njust wondering why is (2^5)-1 here? ",
    "327598": "Congratulations on the win.\n\nThe paper 'A View from the Trenches' references using a 'feature importance weight' to scale down the learning from the positive results, so the network output didn't have a bias toward positive samples, did you implement something similar? \n\nWouldn't a model trained on oversampled positive data predict more positives? ",
    "327205": "Maybe the reason that why removing all raw categorical features works is that your embedding features have captured more information than raw features. Thus raw feature become noise.",
    "327199": "Congratulations.Learn more from your solution. Especially categorical features processing method ",
    "327070": "Congrats! The methods are brilliant! Your methods and features are quite different from us, which is really amazing. I am sure I can learn a lot from you. Thanks!",
    "327036": "Beautiful solution! Thank you very much for sharing and congrats! ",
    "326935": "hello，I'm a beginner in kaggle competitions. can you explain how to  created a bunch of features in a brute-force way?\nThanks",
    "326815": "Thanks for sharing guys. I admire the efficiency of the downsampling for training, yet the breadth of feature exploration. Nice work!",
    "326743": "Thanks and congrats!  \nAnd just a quick question: when you use subsample to decrease the size of negative labels, do you select them randomly every time you train model or just use the same sets every time? ",
    "326732": "I'm a beginner in kaggle competitions.I‘m wondered how to bag LGB model.Does it use the output of last layer feature transform as the input of current layer?   ",
    "326715": "congtats！awesome work！",
    "326710": "Congratulations!  I have a question:your features are calculate on the full data or on the down-sample features? i mean your models diversity is not only train data difference but also feature differenct or just train data difference?",
    "326708": "Awesome! Your score with downsampling made me happy :)",
    "326697": "Congrats and thx for sharing, these two papers are really nice.  A question : brute-force way will bring many features, so how do you choose useful feature?",
    "326694": "Congratulations! I really like the idea with class balancing - I tried it in other way round  boosting positive examples.  ",
    "326685": "Congrats on the result!  Thanks for sharing.  I am not surprised that relying heavily on matrix factorizations helped a lot.",
    "627607": "",
    "330050": "",
    "1071619": "Thank you for sharing your approach 😃",
    "436248": "thank you",
    "401111": "Thank you for sharing. I will try it.",
    "332137": "Thanks for sharing!",
    "330726": "Congratulations!Thanks for sharing.",
    "328344": "Thank you for sharing! "
  }
}