{
  "id": 56328,
  "title": "[2nd Place Solution] from PPP",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56328",
  "author_name": "Feiyang Pan",
  "post_date": "2018-05-08T16:43:27.661000",
  "votes": 114,
  "comment_count": 60,
  "views": 0,
  "content": "<p>Congrats to the top teams and thanks Kaggle and TalkingData for hosting such a perfect competition. Also congrats to Plantsgo, the new Kaggle grandmaster! </p>\n\n<p>Overall the competition was wonderful (except for D**k's Kernel) and we learned much during the last month. Here I'd like to briefly describe our solution as well as some important techniques we used. To summarize, our solution consists of: </p>\n\n<ul>\n<li><p>a framework that is both time and memory efficient to cope with the large dataset,</p></li>\n<li><p>some regular features,</p></li>\n<li><p>two methods, LightGBM and NN,</p></li>\n<li><p>a simple weighted average ensemble of predictions.</p></li>\n</ul>\n\n<h2>Framework</h2>\n\n<p>It is essential for us to use sub-sampling to reduce time and memory costs. At the early stage, we found it so hard to deal with the whole dataset due to the limitation of RAM (128G for Plantsgo, 64G for me, and 32G for Piupiu). Piupiu bought another 16G RAM immediately, but he was upset when finding that it merely helps. The difficulties to deal with such large data are 2-folds: it's hard to extract features, and it's slow to train a model. As the training data is extremely imbalanced, we reduced the data size by sampling a small fraction (5%) of negative samples. In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data. In the training phase, only the sub-sampled training samples were used so it was about 10+ times faster than directly training the entire data. We used 5-fold CV to see the offline performances. It took about half an hour to train one fold (with an Intel i7 core and two hundred of features).</p>\n\n<h2>Feature engineering</h2>\n\n<p>Our team did not have any magic features, although my teammates Plantsgo and Piupiu are both feature engineering experts. Our features were regular in the sense that almost all our features were open-sourced by others in the Kernels one or two weeks after we had used them. What a sad story. There were several kind of features: count features, cumcount features, time-delta features, unique-count features, and which [app/os/channel]s each IP appears in the data. Because we have the sub-sampling framework, we have enough memory space to get hundreds of features.</p>\n\n<h2>Models</h2>\n\n<p>We used two methods, LightGBM and NN. The best single model was LightGBM with Plantsgo's features which scored 0.9837 on the private LB. I have been trying to make a strong neural network to beat LightGBM during the whole month, but obviously I failed. Our best NN scored 0.9834 on the private LB which had a dot-product layer for categorical inputs and deep fully-connected layers for continuous numerical inputs. I believe there must be better NN structures and I really hope to learn it from other top teams. </p>\n\n<h2>Ensemble</h2>\n\n<p>The three of us had three LGB predictions and three NN predictions. So we averaged these 6 predictions by trivially applying some weights inferred from their public LB scores. Now that the private scores are revealed, we find it more correlated to the offline CV scores rather than the public scores. If we had trusted the offline scores, our final score could have been better. </p>\n\n<h3>The leak on the test set: a sad story</h3>\n\n<p>In a word, we did not use the leak on test set, even though it was us to post the topic to Discussion. We were filled with grief when we heard that the leak would help improve about 0.0004. </p>\n\n<p>Thanks for reading it! </p>\n\n<p>也谢谢大家的支持！</p>",
  "messages": [
    {
      "id": 325630,
      "postDate": "2018-05-08T16:43:27.660Z",
      "content": "<p>Congrats to the top teams and thanks Kaggle and TalkingData for hosting such a perfect competition. Also congrats to Plantsgo, the new Kaggle grandmaster! </p>\n\n<p>Overall the competition was wonderful (except for D**k's Kernel) and we learned much during the last month. Here I'd like to briefly describe our solution as well as some important techniques we used. To summarize, our solution consists of: </p>\n\n<ul>\n<li><p>a framework that is both time and memory efficient to cope with the large dataset,</p></li>\n<li><p>some regular features,</p></li>\n<li><p>two methods, LightGBM and NN,</p></li>\n<li><p>a simple weighted average ensemble of predictions.</p></li>\n</ul>\n\n<h2>Framework</h2>\n\n<p>It is essential for us to use sub-sampling to reduce time and memory costs. At the early stage, we found it so hard to deal with the whole dataset due to the limitation of RAM (128G for Plantsgo, 64G for me, and 32G for Piupiu). Piupiu bought another 16G RAM immediately, but he was upset when finding that it merely helps. The difficulties to deal with such large data are 2-folds: it's hard to extract features, and it's slow to train a model. As the training data is extremely imbalanced, we reduced the data size by sampling a small fraction (5%) of negative samples. In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data. In the training phase, only the sub-sampled training samples were used so it was about 10+ times faster than directly training the entire data. We used 5-fold CV to see the offline performances. It took about half an hour to train one fold (with an Intel i7 core and two hundred of features).</p>\n\n<h2>Feature engineering</h2>\n\n<p>Our team did not have any magic features, although my teammates Plantsgo and Piupiu are both feature engineering experts. Our features were regular in the sense that almost all our features were open-sourced by others in the Kernels one or two weeks after we had used them. What a sad story. There were several kind of features: count features, cumcount features, time-delta features, unique-count features, and which [app/os/channel]s each IP appears in the data. Because we have the sub-sampling framework, we have enough memory space to get hundreds of features.</p>\n\n<h2>Models</h2>\n\n<p>We used two methods, LightGBM and NN. The best single model was LightGBM with Plantsgo's features which scored 0.9837 on the private LB. I have been trying to make a strong neural network to beat LightGBM during the whole month, but obviously I failed. Our best NN scored 0.9834 on the private LB which had a dot-product layer for categorical inputs and deep fully-connected layers for continuous numerical inputs. I believe there must be better NN structures and I really hope to learn it from other top teams. </p>\n\n<h2>Ensemble</h2>\n\n<p>The three of us had three LGB predictions and three NN predictions. So we averaged these 6 predictions by trivially applying some weights inferred from their public LB scores. Now that the private scores are revealed, we find it more correlated to the offline CV scores rather than the public scores. If we had trusted the offline scores, our final score could have been better. </p>\n\n<h3>The leak on the test set: a sad story</h3>\n\n<p>In a word, we did not use the leak on test set, even though it was us to post the topic to Discussion. We were filled with grief when we heard that the leak would help improve about 0.0004. </p>\n\n<p>Thanks for reading it! </p>\n\n<p>也谢谢大家的支持！</p>",
      "rawMarkdown": "Congrats to the top teams and thanks Kaggle and TalkingData for hosting such a perfect competition. Also congrats to Plantsgo, the new Kaggle grandmaster! \n\nOverall the competition was wonderful (except for D\\*\\*k's Kernel) and we learned much during the last month. Here I'd like to briefly describe our solution as well as some important techniques we used. To summarize, our solution consists of: \n\n- a framework that is both time and memory efficient to cope with the large dataset,\n\n- some regular features,\n\n- two methods, LightGBM and NN,\n\n- a simple weighted average ensemble of predictions.\n\n## Framework\nIt is essential for us to use sub-sampling to reduce time and memory costs. At the early stage, we found it so hard to deal with the whole dataset due to the limitation of RAM (128G for Plantsgo, 64G for me, and 32G for Piupiu). Piupiu bought another 16G RAM immediately, but he was upset when finding that it merely helps. The difficulties to deal with such large data are 2-folds: it's hard to extract features, and it's slow to train a model. As the training data is extremely imbalanced, we reduced the data size by sampling a small fraction (5%) of negative samples. In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data. In the training phase, only the sub-sampled training samples were used so it was about 10+ times faster than directly training the entire data. We used 5-fold CV to see the offline performances. It took about half an hour to train one fold (with an Intel i7 core and two hundred of features).\n\n## Feature engineering\nOur team did not have any magic features, although my teammates Plantsgo and Piupiu are both feature engineering experts. Our features were regular in the sense that almost all our features were open-sourced by others in the Kernels one or two weeks after we had used them. What a sad story. There were several kind of features: count features, cumcount features, time-delta features, unique-count features, and which [app/os/channel]s each IP appears in the data. Because we have the sub-sampling framework, we have enough memory space to get hundreds of features.\n\n## Models\nWe used two methods, LightGBM and NN. The best single model was LightGBM with Plantsgo's features which scored 0.9837 on the private LB. I have been trying to make a strong neural network to beat LightGBM during the whole month, but obviously I failed. Our best NN scored 0.9834 on the private LB which had a dot-product layer for categorical inputs and deep fully-connected layers for continuous numerical inputs. I believe there must be better NN structures and I really hope to learn it from other top teams. \n\n## Ensemble\nThe three of us had three LGB predictions and three NN predictions. So we averaged these 6 predictions by trivially applying some weights inferred from their public LB scores. Now that the private scores are revealed, we find it more correlated to the offline CV scores rather than the public scores. If we had trusted the offline scores, our final score could have been better. \n\n### The leak on the test set: a sad story\nIn a word, we did not use the leak on test set, even though it was us to post the topic to Discussion. We were filled with grief when we heard that the leak would help improve about 0.0004. \n\nThanks for reading it! \n\n也谢谢大家的支持！",
      "votes": 113
    },
    {
      "id": 325674,
      "postDate": "2018-05-08T17:47:30.063Z",
      "content": "<p>Congrats on the result and the approach!</p>\n\n<p>Is your dot product level similar to libfm or libffm?  I mean, do you have one embedding per categorical that you reuse for all pair interactions, or do you have one embedding per pair of interactions?</p>",
      "rawMarkdown": "Congrats on the result and the approach!\n\nIs your dot product level similar to libfm or libffm?  I mean, do you have one embedding per categorical that you reuse for all pair interactions, or do you have one embedding per pair of interactions?",
      "votes": 3,
      "replies": [
        {
          "id": 325699,
          "postDate": "2018-05-08T18:29:32.840Z",
          "content": "<p>Thanks. It's similar to FM, one embedding per categorical feature. I did not implement an FFM-like structure considering that the memory usage on GPU may be too large.</p>",
          "rawMarkdown": "Thanks. It's similar to FM, one embedding per categorical feature. I did not implement an FFM-like structure considering that the memory usage on GPU may be too large.",
          "votes": 4
        }
      ]
    },
    {
      "id": 325999,
      "postDate": "2018-05-09T06:45:22.110Z",
      "content": "<p>FeiYang\nThanks for sharing and Congrats! 大佬天秀！</p>",
      "rawMarkdown": "FeiYang\nThanks for sharing and Congrats! 大佬天秀！",
      "votes": 1
    },
    {
      "id": 325820,
      "postDate": "2018-05-08T22:54:43.090Z",
      "content": "<p>Brilliant strategies and congrats!</p>",
      "rawMarkdown": "Brilliant strategies and congrats!",
      "votes": 1
    },
    {
      "id": 325665,
      "postDate": "2018-05-08T17:23:20.053Z",
      "content": "<p>Congrats!!!</p>",
      "rawMarkdown": "Congrats!!!",
      "votes": 1
    },
    {
      "id": 326126,
      "postDate": "2018-05-09T10:19:33.007Z",
      "content": "<p>Congrats，Da Lao，I am very interested in your NN model，can you share your NN model details and your scheme about \n how to design your NN model？Really thanks。</p>",
      "rawMarkdown": "Congrats，Da Lao，I am very interested in your NN model，can you share your NN model details and your scheme about \n how to design your NN model？Really thanks。",
      "votes": 2,
      "replies": [
        {
          "id": 326138,
          "postDate": "2018-05-09T10:44:14.693Z",
          "content": "<p>Categorical inputs are embedded and fed into an FM-like dot-product layer. Numerical inputs are fed into 3 FC layers. Then we concatenate the outputs and fed to the last FC layer. </p>",
          "rawMarkdown": "Categorical inputs are embedded and fed into an FM-like dot-product layer. Numerical inputs are fed into 3 FC layers. Then we concatenate the outputs and fed to the last FC layer. ",
          "votes": 4
        }
      ]
    },
    {
      "id": 325890,
      "postDate": "2018-05-09T02:44:16.760Z",
      "content": "<p>Congrats and thanks for sharing! I have some question about it.</p>\n\n<ol>\n<li>Could you talk about the leak on the test set? What is the leak? How do you find it?</li>\n<li>Would you plan to share your complete project code on Kaggle or Github?</li>\n</ol>\n\n<p>Thank a lot!</p>",
      "rawMarkdown": "Congrats and thanks for sharing! I have some question about it.\n\n1. Could you talk about the leak on the test set? What is the leak? How do you find it?\n2. Would you plan to share your complete project code on Kaggle or Github?\n\nThank a lot!",
      "votes": 2,
      "replies": [
        {
          "id": 325980,
          "postDate": "2018-05-09T06:32:36.703Z",
          "content": "<p>The leak is about duplicated samples with different labels in both train and test set, which was first shared <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\">here</a> by our teammate Plantsgo. Some of the top teams found an amazing way to use it which helps improve about 0.0005, have a look at <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">AhmetErdem's solution</a> and <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">CPMP's solution</a>. </p>\n\n<p>The codes are messy so we probably will not make it public.</p>",
          "rawMarkdown": "The leak is about duplicated samples with different labels in both train and test set, which was first shared [here][1] by our teammate Plantsgo. Some of the top teams found an amazing way to use it which helps improve about 0.0005, have a look at [AhmetErdem's solution][2] and [CPMP's solution][3]. \n\nThe codes are messy so we probably will not make it public.\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677 \"Plantsgo's share about the leak\"\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268 \"AhmetErdem's solution\"\n  [3]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283 \"CPMP's solution\"",
          "votes": 2
        },
        {
          "id": 326020,
          "postDate": "2018-05-09T07:02:13.913Z",
          "content": "<p>Thanks for answering. Cool job!</p>",
          "rawMarkdown": "Thanks for answering. Cool job!",
          "votes": 1
        }
      ]
    },
    {
      "id": 325872,
      "postDate": "2018-05-09T01:59:16.937Z",
      "content": "<p>Congratulation! Thanks for sharing, Da Lao!</p>\n\n<p>May I ask several questions here ~~~\n1.  As you mentioned above, your team use 5-fold CV to see the offline performances. Does that mean you also use cv to predict online test, or just to do feature selection\n2. It seems no history target encoding features in your Feature engineering part. I personally spent lots of time creating different group-by target encoding features, and I want to understand how you capture this information (just use the raw category features?)\n3. Did you spend lots of time on feature selection, like forward and backward feature selection ~~ </p>\n\n<p>灰常感谢！！</p>",
      "rawMarkdown": "Congratulation! Thanks for sharing, Da Lao!\n\nMay I ask several questions here ~~~\n1.  As you mentioned above, your team use 5-fold CV to see the offline performances. Does that mean you also use cv to predict online test, or just to do feature selection\n2. It seems no history target encoding features in your Feature engineering part. I personally spent lots of time creating different group-by target encoding features, and I want to understand how you capture this information (just use the raw category features?)\n3. Did you spend lots of time on feature selection, like forward and backward feature selection ~~ \n\n灰常感谢！！\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 325987,
          "postDate": "2018-05-09T06:39:16.883Z",
          "content": "<ol>\n<li><p>The 5-fold CV is also used to make predictions. </p></li>\n<li><p>We did not use target encoding features, because we were afraid that target encoding would somehow overfits the training data.</p></li>\n<li><p>Yes, almost every day we spent a lot of time on selecting features.</p></li>\n</ol>",
          "rawMarkdown": "1. The 5-fold CV is also used to make predictions. \n\n2. We did not use target encoding features, because we were afraid that target encoding would somehow overfits the training data.\n\n3. Yes, almost every day we spent a lot of time on selecting features.",
          "votes": 2
        }
      ]
    },
    {
      "id": 331897,
      "postDate": "2018-05-22T04:22:36.430Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 329217,
      "postDate": "2018-05-16T02:59:13.110Z",
      "content": "<p>When you made a submission,the csv was generated from sample dataset or the entire dataset?</p>",
      "rawMarkdown": "When you made a submission,the csv was generated from sample dataset or the entire dataset?"
    },
    {
      "id": 328908,
      "postDate": "2018-05-15T10:54:31.080Z",
      "content": "<p>Nice solution and It was very interesting to compete with your teams.\nbtw, I'm just curious about how much the score will become better if we combine the best submission.\ncould you upload your best submission if possible? \n1st &amp; 5th &amp; 6th has already uploaded.\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423</a></p>",
      "rawMarkdown": "Nice solution and It was very interesting to compete with your teams.\nbtw, I'm just curious about how much the score will become better if we combine the best submission.\ncould you upload your best submission if possible? \n1st &amp; 5th &amp; 6th has already uploaded.\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423"
    },
    {
      "id": 326651,
      "postDate": "2018-05-10T05:19:33.913Z",
      "content": "<p>congrats, i also implemented a NN like fm (actually, refering to DeepFM) , but it turned out to be a bad result. Can dalao share your NN code or the structure in github? </p>",
      "rawMarkdown": "congrats, i also implemented a NN like fm (actually, refering to DeepFM) , but it turned out to be a bad result. Can dalao share your NN code or the structure in github? "
    },
    {
      "id": 326635,
      "postDate": "2018-05-10T04:53:43.040Z",
      "content": "<p>Congrats，Da Lao，I have some question about  sub-sampling,  I want to know, the sub-sampling is sampled from the train or the all data, like the ijcai competition，I will user the 20 21 22 to train, and to predict the 23th. do you mean you sampled the sub-sampling from 20 21 22 ?</p>",
      "rawMarkdown": "Congrats，Da Lao，I have some question about  sub-sampling,  I want to know, the sub-sampling is sampled from the train or the all data, like the ijcai competition，I will user the 20 21 22 to train, and to predict the 23th. do you mean you sampled the sub-sampling from 20 21 22 ?"
    },
    {
      "id": 326454,
      "postDate": "2018-05-09T18:50:04.917Z",
      "content": "<p>Thanks for sharing! Take out my small notebook and write down all details.....</p>",
      "rawMarkdown": "Thanks for sharing! Take out my small notebook and write down all details....."
    },
    {
      "id": 326219,
      "postDate": "2018-05-09T13:04:50.693Z",
      "content": "<p>Da  Lao，when you use sample dataset，how did you setup your lightgbm hyperparameter？what is the value of scale_pos_weight？ How did you deal with sample dataset imbalance？</p>",
      "rawMarkdown": "Da  Lao，when you use sample dataset，how did you setup your lightgbm hyperparameter？what is the value of scale_pos_weight？ How did you deal with sample dataset imbalance？",
      "replies": [
        {
          "id": 326235,
          "postDate": "2018-05-09T13:21:39.793Z",
          "content": "<p>Good question. Intuitively, the scale_pos_weight param is directly proportional to the ratio of negative samples. So if you set scale_pos_weight to 100 when training the entire dataset, you should set it to 100*5%=5 after sampling 5% negative samples. (You can also tune the parameter, it does not make a big difference.)</p>",
          "rawMarkdown": "Good question. Intuitively, the scale_pos_weight param is directly proportional to the ratio of negative samples. So if you set scale_pos_weight to 100 when training the entire dataset, you should set it to 100*5%=5 after sampling 5% negative samples. (You can also tune the parameter, it does not make a big difference.)",
          "votes": 2
        }
      ]
    },
    {
      "id": 326098,
      "postDate": "2018-05-09T09:19:39.373Z",
      "content": "<p>Thx for sharing and congrats!</p>",
      "rawMarkdown": "Thx for sharing and congrats!"
    },
    {
      "id": 326096,
      "postDate": "2018-05-09T09:10:53.557Z",
      "content": "<p>Thanks for sharing and Congrats!!! One question: Is your best single model training on the sub-sample data? Or you only use sub-sample data to select features and then train lgb on all data?</p>",
      "rawMarkdown": "Thanks for sharing and Congrats!!! One question: Is your best single model training on the sub-sample data? Or you only use sub-sample data to select features and then train lgb on all data?",
      "replies": [
        {
          "id": 326129,
          "postDate": "2018-05-09T10:29:33.587Z",
          "content": "<p>The best single model is trained on the sub-sampled data and the submission is the average prediction of 5-fold CV.</p>",
          "rawMarkdown": "The best single model is trained on the sub-sampled data and the submission is the average prediction of 5-fold CV.",
          "votes": 1
        },
        {
          "id": 326157,
          "postDate": "2018-05-09T11:15:54.930Z",
          "content": "<p>Thanks for your reply</p>",
          "rawMarkdown": "Thanks for your reply"
        }
      ]
    },
    {
      "id": 326070,
      "postDate": "2018-05-09T08:32:21.857Z",
      "content": "<p>some question about \"which [app/os/channel]s each IP appears in the data\", does this mean you cal each ip's click counts on every app /os/channle?is so ,there are so many app / os /channles，how do you solve this problem?</p>",
      "rawMarkdown": "some question about \"which [app/os/channel]s each IP appears in the data\", does this mean you cal each ip's click counts on every app /os/channle?is so ,there are so many app / os /channles，how do you solve this problem?",
      "replies": [
        {
          "id": 326136,
          "postDate": "2018-05-09T10:40:08.250Z",
          "content": "<p>Yes. We only picked several most frequent app/os/channels.</p>",
          "rawMarkdown": "Yes. We only picked several most frequent app/os/channels.",
          "votes": 1
        },
        {
          "id": 326228,
          "postDate": "2018-05-09T13:17:40.830Z",
          "content": "<p>thank you for your reply, like @CMPP 's  method，he uses SVD to solve this problem（多谢）</p>",
          "rawMarkdown": "thank you for your reply, like @CMPP 's  method，he uses SVD to solve this problem（多谢）"
        }
      ]
    },
    {
      "id": 326064,
      "postDate": "2018-05-09T08:12:00.173Z",
      "content": "<p>大佬比赛的胜率这么高，膜拜一下</p>",
      "rawMarkdown": "大佬比赛的胜率这么高，膜拜一下",
      "replies": [
        {
          "id": 326133,
          "postDate": "2018-05-09T10:37:39.797Z",
          "content": "<p>I am very proud of our team, PPP. Plantsgo and Piupiu are great teammates and excellent competitors. We all have top-class solo performances and even better scores as a team. </p>",
          "rawMarkdown": "I am very proud of our team, PPP. Plantsgo and Piupiu are great teammates and excellent competitors. We all have top-class solo performances and even better scores as a team. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 325954,
      "postDate": "2018-05-09T05:36:18.123Z",
      "content": "<p>Congratulation PPP team!</p>\n\n<blockquote>\n  <p>We used 5-fold CV to see the offline performances.\n  Can you explain why do you choose 5-fold CV and how do you use your the five scores in the competition? </p>\n</blockquote>",
      "rawMarkdown": "Congratulation PPP team!\n\n&gt;  We used 5-fold CV to see the offline performances.\nCan you explain why do you choose 5-fold CV and how do you use your the five scores in the competition? "
    },
    {
      "id": 325909,
      "postDate": "2018-05-09T03:29:29.747Z",
      "content": "<p>Brilliant and Congrats! I really hope to learn more from you and your team in the coming years. You guys are truly Data Guru!~</p>",
      "rawMarkdown": "Brilliant and Congrats! I really hope to learn more from you and your team in the coming years. You guys are truly Data Guru!~"
    },
    {
      "id": 325900,
      "postDate": "2018-05-09T03:11:54.073Z",
      "content": "<p>Thanks for sharing! train phase only subsampling is so briliant! 感谢大佬！</p>",
      "rawMarkdown": "Thanks for sharing! train phase only subsampling is so briliant! 感谢大佬！"
    },
    {
      "id": 325888,
      "postDate": "2018-05-09T02:34:09.910Z",
      "content": "<p>Congrats and thanks for sharing!  Maybey someday we can have a cup of tea together~~~</p>",
      "rawMarkdown": "Congrats and thanks for sharing!  Maybey someday we can have a cup of tea together~~~"
    },
    {
      "id": 325882,
      "postDate": "2018-05-09T02:21:44.640Z",
      "content": "<p>Congrats! 大佬天秀！</p>",
      "rawMarkdown": "Congrats! 大佬天秀！"
    },
    {
      "id": 325866,
      "postDate": "2018-05-09T01:36:42.347Z",
      "content": "<p>Niubi~~~ May I ask one question, how did you guys do 5-fold cv with a time-based dataset?Thanks~~~</p>",
      "rawMarkdown": "Niubi~~~ May I ask one question, how did you guys do 5-fold cv with a time-based dataset?Thanks~~~",
      "replies": [
        {
          "id": 325995,
          "postDate": "2018-05-09T06:43:33.380Z",
          "content": "<p>It is the trivial 5-fold CV we used. We do not guarantee the correctness of K-fold CV if the features and samples involve timestamps, but in Kaggle contests we can try and see if it works.</p>",
          "rawMarkdown": "It is the trivial 5-fold CV we used. We do not guarantee the correctness of K-fold CV if the features and samples involve timestamps, but in Kaggle contests we can try and see if it works.",
          "votes": 2
        },
        {
          "id": 326022,
          "postDate": "2018-05-09T07:03:19.577Z",
          "content": "<p>Thanks,laotie!</p>",
          "rawMarkdown": "Thanks,laotie!"
        },
        {
          "id": 326108,
          "postDate": "2018-05-09T09:43:41.083Z",
          "content": "<p>what does it mean by trivial 5-fold cv... i  failed to search it in google...</p>",
          "rawMarkdown": "what does it mean by trivial 5-fold cv... i  failed to search it in google...",
          "votes": 1
        },
        {
          "id": 326156,
          "postDate": "2018-05-09T11:11:13.920Z",
          "content": "<p>就是最普通的五折CV。。</p>",
          "rawMarkdown": "就是最普通的五折CV。。",
          "votes": 2
        }
      ]
    },
    {
      "id": 325853,
      "postDate": "2018-05-09T00:53:25.987Z",
      "content": "<p>Congrats and thanks for you share！\nAwesome work！</p>",
      "rawMarkdown": "Congrats and thanks for you share！\nAwesome work！"
    },
    {
      "id": 325819,
      "postDate": "2018-05-08T22:53:58.060Z",
      "content": "<blockquote>\n  <p>In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data.</p>\n</blockquote>\n\n<p>Does this mean train + test + supplement, or just train?</p>",
      "rawMarkdown": "&gt; In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data.\n\nDoes this mean train + test + supplement, or just train?",
      "replies": [
        {
          "id": 326139,
          "postDate": "2018-05-09T10:47:04.763Z",
          "content": "<p>Train + old_test (supplement).</p>",
          "rawMarkdown": "Train + old_test (supplement)."
        }
      ]
    },
    {
      "id": 325798,
      "postDate": "2018-05-08T21:48:24.303Z",
      "content": "<p>Congrats and thanks for sharing! Your tricks are very impressive!</p>",
      "rawMarkdown": "Congrats and thanks for sharing! Your tricks are very impressive!"
    },
    {
      "id": 325794,
      "postDate": "2018-05-08T21:36:32.150Z",
      "content": "<blockquote>\n  <p>we reduced the data size by sampling a small fraction (5%) of negative samples</p>\n</blockquote>\n\n<p>Thanks for sharing! Was this a random sample over all negative examples or was it biased in some way? It seems like next_click features would be affected by this.</p>",
      "rawMarkdown": "&gt; we reduced the data size by sampling a small fraction (5%) of negative samples\n\nThanks for sharing! Was this a random sample over all negative examples or was it biased in some way? It seems like next_click features would be affected by this.",
      "replies": [
        {
          "id": 325795,
          "postDate": "2018-05-08T21:38:39.910Z",
          "content": "<p>If I understand well, the feature engineering is done once on the whole set. Then the sampling happens.</p>",
          "rawMarkdown": "If I understand well, the feature engineering is done once on the whole set. Then the sampling happens.",
          "votes": 1
        },
        {
          "id": 325813,
          "postDate": "2018-05-08T22:26:20.110Z",
          "content": "<p>For each feature, it is firstly extracted from the entire dataset (as a very large column), then merged into the small sub-sampled data (so only about 5% entries of that column are kept). So the resulted features have correct values. </p>",
          "rawMarkdown": "For each feature, it is firstly extracted from the entire dataset (as a very large column), then merged into the small sub-sampled data (so only about 5% entries of that column are kept). So the resulted features have correct values. ",
          "votes": 6
        }
      ]
    },
    {
      "id": 325760,
      "postDate": "2018-05-08T20:41:40.487Z",
      "content": "<p>Thanks for sharing! I desperately tried to find a magic feature and had very high hopes with features extracted out of performing FFT on rolling windows but it turned out MUCH too computationally expensive to be computed in a timely manner.\nCongratulations to you three!</p>",
      "rawMarkdown": "Thanks for sharing! I desperately tried to find a magic feature and had very high hopes with features extracted out of performing FFT on rolling windows but it turned out MUCH too computationally expensive to be computed in a timely manner.\nCongratulations to you three!",
      "replies": [
        {
          "id": 326036,
          "postDate": "2018-05-09T07:12:46.833Z",
          "content": "<p>I also tried some features based on sliding windows but nothing works. :(</p>",
          "rawMarkdown": "I also tried some features based on sliding windows but nothing works. :("
        }
      ]
    },
    {
      "id": 325750,
      "postDate": "2018-05-08T20:11:28.457Z",
      "content": "<p>Thanks for sharing your solution and congrats!</p>",
      "rawMarkdown": "Thanks for sharing your solution and congrats!"
    },
    {
      "id": 325741,
      "postDate": "2018-05-08T19:54:44.863Z",
      "content": "<p>Congrats and thanks for you share,it is amazing.</p>",
      "rawMarkdown": "Congrats and thanks for you share,it is amazing."
    },
    {
      "id": 325659,
      "postDate": "2018-05-08T17:10:27.660Z",
      "content": "<p>Congratulation. It is a really brilliant idea of data sampling to overcome the limit on RAM</p>",
      "rawMarkdown": "Congratulation. It is a really brilliant idea of data sampling to overcome the limit on RAM"
    },
    {
      "id": 325657,
      "postDate": "2018-05-08T17:06:52.810Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 325653,
      "postDate": "2018-05-08T17:03:44.693Z",
      "content": "<p>Brilliant</p>",
      "rawMarkdown": "Brilliant"
    },
    {
      "id": 325655,
      "postDate": "2018-05-08T17:05:43.680Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 325960,
      "postDate": "2018-05-09T05:55:01.433Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 325651,
      "postDate": "2018-05-08T17:02:43.500Z",
      "content": "<p>Thanks for sharing, Da Lao!</p>",
      "rawMarkdown": "Thanks for sharing, Da Lao!",
      "votes": 1
    },
    {
      "id": 330295,
      "postDate": "2018-05-18T13:56:57.943Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!"
    },
    {
      "id": 325933,
      "postDate": "2018-05-09T04:55:48.617Z",
      "content": "<p>NIUBI NIUBI NIUBI</p>\n\n<p>Thanks for sharing.</p>",
      "rawMarkdown": "NIUBI NIUBI NIUBI\n\nThanks for sharing."
    },
    {
      "id": 325857,
      "postDate": "2018-05-09T01:05:31.533Z",
      "content": "<p>Thanks, Da Lao!!!</p>",
      "rawMarkdown": "Thanks, Da Lao!!!"
    },
    {
      "id": 325727,
      "postDate": "2018-05-08T19:44:18.073Z",
      "content": "<p>Thanks for your share! </p>",
      "rawMarkdown": "Thanks for your share! "
    }
  ],
  "comments": [
    {
      "id": 325674,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-08T17:47:30.063000",
      "content": "<p>Congrats on the result and the approach!</p>\n\n<p>Is your dot product level similar to libfm or libffm?  I mean, do you have one embedding per categorical that you reuse for all pair interactions, or do you have one embedding per pair of interactions?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 325699,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-08T18:29:32.840000",
          "content": "<p>Thanks. It's similar to FM, one embedding per categorical feature. I did not implement an FFM-like structure considering that the memory usage on GPU may be too large.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 325999,
      "author_name": "Eric",
      "author_url": "",
      "post_date": "2018-05-09T06:45:22.110000",
      "content": "<p>FeiYang\nThanks for sharing and Congrats! 大佬天秀！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325820,
      "author_name": "Steinhafen",
      "author_url": "",
      "post_date": "2018-05-08T22:54:43.090000",
      "content": "<p>Brilliant strategies and congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325665,
      "author_name": "seaguII",
      "author_url": "",
      "post_date": "2018-05-08T17:23:20.053000",
      "content": "<p>Congrats!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 326126,
      "author_name": "twoone",
      "author_url": "",
      "post_date": "2018-05-09T10:19:33.007000",
      "content": "<p>Congrats，Da Lao，I am very interested in your NN model，can you share your NN model details and your scheme about \n how to design your NN model？Really thanks。</p>",
      "votes": 2,
      "replies": [
        {
          "id": 326138,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T10:44:14.693000",
          "content": "<p>Categorical inputs are embedded and fed into an FM-like dot-product layer. Numerical inputs are fed into 3 FC layers. Then we concatenate the outputs and fed to the last FC layer. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 325890,
      "author_name": "Shawn Xiao",
      "author_url": "",
      "post_date": "2018-05-09T02:44:16.760000",
      "content": "<p>Congrats and thanks for sharing! I have some question about it.</p>\n\n<ol>\n<li>Could you talk about the leak on the test set? What is the leak? How do you find it?</li>\n<li>Would you plan to share your complete project code on Kaggle or Github?</li>\n</ol>\n\n<p>Thank a lot!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 325980,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T06:32:36.703000",
          "content": "<p>The leak is about duplicated samples with different labels in both train and test set, which was first shared <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677\">here</a> by our teammate Plantsgo. Some of the top teams found an amazing way to use it which helps improve about 0.0005, have a look at <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56268\">AhmetErdem's solution</a> and <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">CPMP's solution</a>. </p>\n\n<p>The codes are messy so we probably will not make it public.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326020,
          "author_name": "Shawn Xiao",
          "author_url": "",
          "post_date": "2018-05-09T07:02:13.913000",
          "content": "<p>Thanks for answering. Cool job!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325872,
      "author_name": "fortune cookie",
      "author_url": "",
      "post_date": "2018-05-09T01:59:16.937000",
      "content": "<p>Congratulation! Thanks for sharing, Da Lao!</p>\n\n<p>May I ask several questions here ~~~\n1.  As you mentioned above, your team use 5-fold CV to see the offline performances. Does that mean you also use cv to predict online test, or just to do feature selection\n2. It seems no history target encoding features in your Feature engineering part. I personally spent lots of time creating different group-by target encoding features, and I want to understand how you capture this information (just use the raw category features?)\n3. Did you spend lots of time on feature selection, like forward and backward feature selection ~~ </p>\n\n<p>灰常感谢！！</p>",
      "votes": 2,
      "replies": [
        {
          "id": 325987,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T06:39:16.883000",
          "content": "<ol>\n<li><p>The 5-fold CV is also used to make predictions. </p></li>\n<li><p>We did not use target encoding features, because we were afraid that target encoding would somehow overfits the training data.</p></li>\n<li><p>Yes, almost every day we spent a lot of time on selecting features.</p></li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 331897,
      "author_name": "koooimo",
      "author_url": "",
      "post_date": "2018-05-22T04:22:36.430000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329217,
      "author_name": "twoone",
      "author_url": "",
      "post_date": "2018-05-16T02:59:13.110000",
      "content": "<p>When you made a submission,the csv was generated from sample dataset or the entire dataset?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328908,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2018-05-15T10:54:31.080000",
      "content": "<p>Nice solution and It was very interesting to compete with your teams.\nbtw, I'm just curious about how much the score will become better if we combine the best submission.\ncould you upload your best submission if possible? \n1st &amp; 5th &amp; 6th has already uploaded.\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326651,
      "author_name": "sparkingarthur",
      "author_url": "",
      "post_date": "2018-05-10T05:19:33.913000",
      "content": "<p>congrats, i also implemented a NN like fm (actually, refering to DeepFM) , but it turned out to be a bad result. Can dalao share your NN code or the structure in github? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326635,
      "author_name": "chuanyueaiqing",
      "author_url": "",
      "post_date": "2018-05-10T04:53:43.040000",
      "content": "<p>Congrats，Da Lao，I have some question about  sub-sampling,  I want to know, the sub-sampling is sampled from the train or the all data, like the ijcai competition，I will user the 20 21 22 to train, and to predict the 23th. do you mean you sampled the sub-sampling from 20 21 22 ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326454,
      "author_name": "Pengyue Wang",
      "author_url": "",
      "post_date": "2018-05-09T18:50:04.917000",
      "content": "<p>Thanks for sharing! Take out my small notebook and write down all details.....</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326219,
      "author_name": "twoone",
      "author_url": "",
      "post_date": "2018-05-09T13:04:50.693000",
      "content": "<p>Da  Lao，when you use sample dataset，how did you setup your lightgbm hyperparameter？what is the value of scale_pos_weight？ How did you deal with sample dataset imbalance？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326235,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T13:21:39.793000",
          "content": "<p>Good question. Intuitively, the scale_pos_weight param is directly proportional to the ratio of negative samples. So if you set scale_pos_weight to 100 when training the entire dataset, you should set it to 100*5%=5 after sampling 5% negative samples. (You can also tune the parameter, it does not make a big difference.)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 326098,
      "author_name": "zr",
      "author_url": "",
      "post_date": "2018-05-09T09:19:39.373000",
      "content": "<p>Thx for sharing and congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326096,
      "author_name": "Yam",
      "author_url": "",
      "post_date": "2018-05-09T09:10:53.557000",
      "content": "<p>Thanks for sharing and Congrats!!! One question: Is your best single model training on the sub-sample data? Or you only use sub-sample data to select features and then train lgb on all data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326129,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T10:29:33.587000",
          "content": "<p>The best single model is trained on the sub-sampled data and the submission is the average prediction of 5-fold CV.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326157,
          "author_name": "Yam",
          "author_url": "",
          "post_date": "2018-05-09T11:15:54.930000",
          "content": "<p>Thanks for your reply</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326070,
      "author_name": "PW",
      "author_url": "",
      "post_date": "2018-05-09T08:32:21.857000",
      "content": "<p>some question about \"which [app/os/channel]s each IP appears in the data\", does this mean you cal each ip's click counts on every app /os/channle?is so ,there are so many app / os /channles，how do you solve this problem?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326136,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T10:40:08.250000",
          "content": "<p>Yes. We only picked several most frequent app/os/channels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326228,
          "author_name": "PW",
          "author_url": "",
          "post_date": "2018-05-09T13:17:40.830000",
          "content": "<p>thank you for your reply, like @CMPP 's  method，he uses SVD to solve this problem（多谢）</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 326064,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "2018-05-09T08:12:00.173000",
      "content": "<p>大佬比赛的胜率这么高，膜拜一下</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326133,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T10:37:39.797000",
          "content": "<p>I am very proud of our team, PPP. Plantsgo and Piupiu are great teammates and excellent competitors. We all have top-class solo performances and even better scores as a team. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 325954,
      "author_name": "shark",
      "author_url": "",
      "post_date": "2018-05-09T05:36:18.123000",
      "content": "<p>Congratulation PPP team!</p>\n\n<blockquote>\n  <p>We used 5-fold CV to see the offline performances.\n  Can you explain why do you choose 5-fold CV and how do you use your the five scores in the competition? </p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325909,
      "author_name": "Gerryl",
      "author_url": "",
      "post_date": "2018-05-09T03:29:29.747000",
      "content": "<p>Brilliant and Congrats! I really hope to learn more from you and your team in the coming years. You guys are truly Data Guru!~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325900,
      "author_name": "Jiazhen Xi",
      "author_url": "",
      "post_date": "2018-05-09T03:11:54.073000",
      "content": "<p>Thanks for sharing! train phase only subsampling is so briliant! 感谢大佬！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325888,
      "author_name": "AlexYung",
      "author_url": "",
      "post_date": "2018-05-09T02:34:09.910000",
      "content": "<p>Congrats and thanks for sharing!  Maybey someday we can have a cup of tea together~~~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325882,
      "author_name": "Perfect Is Shit",
      "author_url": "",
      "post_date": "2018-05-09T02:21:44.640000",
      "content": "<p>Congrats! 大佬天秀！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325866,
      "author_name": "YulinGUO",
      "author_url": "",
      "post_date": "2018-05-09T01:36:42.347000",
      "content": "<p>Niubi~~~ May I ask one question, how did you guys do 5-fold cv with a time-based dataset?Thanks~~~</p>",
      "votes": 0,
      "replies": [
        {
          "id": 325995,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T06:43:33.380000",
          "content": "<p>It is the trivial 5-fold CV we used. We do not guarantee the correctness of K-fold CV if the features and samples involve timestamps, but in Kaggle contests we can try and see if it works.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 326022,
          "author_name": "YulinGUO",
          "author_url": "",
          "post_date": "2018-05-09T07:03:19.577000",
          "content": "<p>Thanks,laotie!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 326108,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-09T09:43:41.083000",
          "content": "<p>what does it mean by trivial 5-fold cv... i  failed to search it in google...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 326156,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T11:11:13.920000",
          "content": "<p>就是最普通的五折CV。。</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 325853,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "2018-05-09T00:53:25.987000",
      "content": "<p>Congrats and thanks for you share！\nAwesome work！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325819,
      "author_name": "MatthewEmery",
      "author_url": "",
      "post_date": "2018-05-08T22:53:58.060000",
      "content": "<blockquote>\n  <p>In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data.</p>\n</blockquote>\n\n<p>Does this mean train + test + supplement, or just train?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326139,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T10:47:04.763000",
          "content": "<p>Train + old_test (supplement).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325798,
      "author_name": "Laevatein",
      "author_url": "",
      "post_date": "2018-05-08T21:48:24.303000",
      "content": "<p>Congrats and thanks for sharing! Your tricks are very impressive!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325794,
      "author_name": "astraldawn",
      "author_url": "",
      "post_date": "2018-05-08T21:36:32.150000",
      "content": "<blockquote>\n  <p>we reduced the data size by sampling a small fraction (5%) of negative samples</p>\n</blockquote>\n\n<p>Thanks for sharing! Was this a random sample over all negative examples or was it biased in some way? It seems like next_click features would be affected by this.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 325795,
          "author_name": "Dany Majard",
          "author_url": "",
          "post_date": "2018-05-08T21:38:39.910000",
          "content": "<p>If I understand well, the feature engineering is done once on the whole set. Then the sampling happens.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325813,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-08T22:26:20.110000",
          "content": "<p>For each feature, it is firstly extracted from the entire dataset (as a very large column), then merged into the small sub-sampled data (so only about 5% entries of that column are kept). So the resulted features have correct values. </p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 325760,
      "author_name": "Dany Majard",
      "author_url": "",
      "post_date": "2018-05-08T20:41:40.487000",
      "content": "<p>Thanks for sharing! I desperately tried to find a magic feature and had very high hopes with features extracted out of performing FFT on rolling windows but it turned out MUCH too computationally expensive to be computed in a timely manner.\nCongratulations to you three!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326036,
          "author_name": "Feiyang Pan",
          "author_url": "",
          "post_date": "2018-05-09T07:12:46.833000",
          "content": "<p>I also tried some features based on sliding windows but nothing works. :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325750,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T20:11:28.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325741,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T19:54:44.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325659,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T17:10:27.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325657,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T17:06:52.810000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325653,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T17:03:44.693000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325655,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T17:05:43.680000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325960,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-09T05:55:01.433000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 325651,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T17:02:43.500000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 330295,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-18T13:56:57.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325933,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-09T04:55:48.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325857,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-09T01:05:31.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325727,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T19:44:18.073000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325630": "Congrats to the top teams and thanks Kaggle and TalkingData for hosting such a perfect competition. Also congrats to Plantsgo, the new Kaggle grandmaster! \n\nOverall the competition was wonderful (except for D\\*\\*k's Kernel) and we learned much during the last month. Here I'd like to briefly describe our solution as well as some important techniques we used. To summarize, our solution consists of: \n\n- a framework that is both time and memory efficient to cope with the large dataset,\n\n- some regular features,\n\n- two methods, LightGBM and NN,\n\n- a simple weighted average ensemble of predictions.\n\n## Framework\nIt is essential for us to use sub-sampling to reduce time and memory costs. At the early stage, we found it so hard to deal with the whole dataset due to the limitation of RAM (128G for Plantsgo, 64G for me, and 32G for Piupiu). Piupiu bought another 16G RAM immediately, but he was upset when finding that it merely helps. The difficulties to deal with such large data are 2-folds: it's hard to extract features, and it's slow to train a model. As the training data is extremely imbalanced, we reduced the data size by sampling a small fraction (5%) of negative samples. In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data. In the training phase, only the sub-sampled training samples were used so it was about 10+ times faster than directly training the entire data. We used 5-fold CV to see the offline performances. It took about half an hour to train one fold (with an Intel i7 core and two hundred of features).\n\n## Feature engineering\nOur team did not have any magic features, although my teammates Plantsgo and Piupiu are both feature engineering experts. Our features were regular in the sense that almost all our features were open-sourced by others in the Kernels one or two weeks after we had used them. What a sad story. There were several kind of features: count features, cumcount features, time-delta features, unique-count features, and which [app/os/channel]s each IP appears in the data. Because we have the sub-sampling framework, we have enough memory space to get hundreds of features.\n\n## Models\nWe used two methods, LightGBM and NN. The best single model was LightGBM with Plantsgo's features which scored 0.9837 on the private LB. I have been trying to make a strong neural network to beat LightGBM during the whole month, but obviously I failed. Our best NN scored 0.9834 on the private LB which had a dot-product layer for categorical inputs and deep fully-connected layers for continuous numerical inputs. I believe there must be better NN structures and I really hope to learn it from other top teams. \n\n## Ensemble\nThe three of us had three LGB predictions and three NN predictions. So we averaged these 6 predictions by trivially applying some weights inferred from their public LB scores. Now that the private scores are revealed, we find it more correlated to the offline CV scores rather than the public scores. If we had trusted the offline scores, our final score could have been better. \n\n### The leak on the test set: a sad story\nIn a word, we did not use the leak on test set, even though it was us to post the topic to Discussion. We were filled with grief when we heard that the leak would help improve about 0.0004. \n\nThanks for reading it! \n\n也谢谢大家的支持！",
    "325674": "Congrats on the result and the approach!\n\nIs your dot product level similar to libfm or libffm?  I mean, do you have one embedding per categorical that you reuse for all pair interactions, or do you have one embedding per pair of interactions?",
    "325999": "FeiYang\nThanks for sharing and Congrats! 大佬天秀！",
    "325820": "Brilliant strategies and congrats!",
    "325665": "Congrats!!!",
    "326126": "Congrats，Da Lao，I am very interested in your NN model，can you share your NN model details and your scheme about \n how to design your NN model？Really thanks。",
    "325890": "Congrats and thanks for sharing! I have some question about it.\n\n1. Could you talk about the leak on the test set? What is the leak? How do you find it?\n2. Would you plan to share your complete project code on Kaggle or Github?\n\nThank a lot!",
    "325872": "Congratulation! Thanks for sharing, Da Lao!\n\nMay I ask several questions here ~~~\n1.  As you mentioned above, your team use 5-fold CV to see the offline performances. Does that mean you also use cv to predict online test, or just to do feature selection\n2. It seems no history target encoding features in your Feature engineering part. I personally spent lots of time creating different group-by target encoding features, and I want to understand how you capture this information (just use the raw category features?)\n3. Did you spend lots of time on feature selection, like forward and backward feature selection ~~ \n\n灰常感谢！！\n\n",
    "331897": "Congrats!",
    "329217": "When you made a submission,the csv was generated from sample dataset or the entire dataset?",
    "328908": "Nice solution and It was very interesting to compete with your teams.\nbtw, I'm just curious about how much the score will become better if we combine the best submission.\ncould you upload your best submission if possible? \n1st &amp; 5th &amp; 6th has already uploaded.\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56423",
    "326651": "congrats, i also implemented a NN like fm (actually, refering to DeepFM) , but it turned out to be a bad result. Can dalao share your NN code or the structure in github? ",
    "326635": "Congrats，Da Lao，I have some question about  sub-sampling,  I want to know, the sub-sampling is sampled from the train or the all data, like the ijcai competition，I will user the 20 21 22 to train, and to predict the 23th. do you mean you sampled the sub-sampling from 20 21 22 ?",
    "326454": "Thanks for sharing! Take out my small notebook and write down all details.....",
    "326219": "Da  Lao，when you use sample dataset，how did you setup your lightgbm hyperparameter？what is the value of scale_pos_weight？ How did you deal with sample dataset imbalance？",
    "326098": "Thx for sharing and congrats!",
    "326096": "Thanks for sharing and Congrats!!! One question: Is your best single model training on the sub-sample data? Or you only use sub-sample data to select features and then train lgb on all data?",
    "326070": "some question about \"which [app/os/channel]s each IP appears in the data\", does this mean you cal each ip's click counts on every app /os/channle?is so ,there are so many app / os /channles，how do you solve this problem?",
    "326064": "大佬比赛的胜率这么高，膜拜一下",
    "325954": "Congratulation PPP team!\n\n&gt;  We used 5-fold CV to see the offline performances.\nCan you explain why do you choose 5-fold CV and how do you use your the five scores in the competition? ",
    "325909": "Brilliant and Congrats! I really hope to learn more from you and your team in the coming years. You guys are truly Data Guru!~",
    "325900": "Thanks for sharing! train phase only subsampling is so briliant! 感谢大佬！",
    "325888": "Congrats and thanks for sharing!  Maybey someday we can have a cup of tea together~~~",
    "325882": "Congrats! 大佬天秀！",
    "325866": "Niubi~~~ May I ask one question, how did you guys do 5-fold cv with a time-based dataset?Thanks~~~",
    "325853": "Congrats and thanks for you share！\nAwesome work！",
    "325819": "&gt; In the feature engineering phase, new features were extracted from the entire data and merged into the sub-sampled data.\n\nDoes this mean train + test + supplement, or just train?",
    "325798": "Congrats and thanks for sharing! Your tricks are very impressive!",
    "325794": "&gt; we reduced the data size by sampling a small fraction (5%) of negative samples\n\nThanks for sharing! Was this a random sample over all negative examples or was it biased in some way? It seems like next_click features would be affected by this.",
    "325760": "Thanks for sharing! I desperately tried to find a magic feature and had very high hopes with features extracted out of performing FFT on rolling windows but it turned out MUCH too computationally expensive to be computed in a timely manner.\nCongratulations to you three!",
    "325750": "Thanks for sharing your solution and congrats!",
    "325741": "Congrats and thanks for you share,it is amazing.",
    "325659": "Congratulation. It is a really brilliant idea of data sampling to overcome the limit on RAM",
    "325657": "Congrats!",
    "325653": "Brilliant",
    "325655": "",
    "325960": "Thanks for sharing",
    "325651": "Thanks for sharing, Da Lao!",
    "330295": "Congrats and thanks for sharing!",
    "325933": "NIUBI NIUBI NIUBI\n\nThanks for sharing.",
    "325857": "Thanks, Da Lao!!!",
    "325727": "Thanks for your share! "
  }
}