{
  "id": 27977,
  "title": "2nd place solution | team brain-afk",
  "url": "/competitions/outbrain-click-prediction/writeups/brain-afk-2nd-place-solution-team-brain-afk",
  "author_name": "",
  "post_date": "2017-01-21T15:53:22.413Z",
  "votes": 27,
  "comment_count": 11,
  "views": 848,
  "content": "<p>Hi fellow kagglers,</p>\n\n<p>Before starting ... a big thanks to Outbrain and Kaggle for hosting this competition.  Also congratulations to code monkeys for a well deserved win, to Three Data Points and last but not least to my teammates: This has been a really fun competition, the teamwork was just great and there was plenty of things to learn.</p>\n\n<p>Our solution involved a lot of feature engineering, model diversity, stacking and a custum implementation of FFM by Alexey, which allowed very fast iterations and yielded a very strong single model at the end.</p>\n\n<p><br><strong>Our most important features</strong><br></p>\n\n<ul>\n<li>rcarson's leak feature. This was very strong. It was slightly\nimproved by    bucketing it into rows where the page_view doc was\nbefore, 1 hour after, 1 day after and &gt; 1 day after the display\ntimestamp.</li>\n<li>Competing ads – on each row, we hashed each individual competing ad\nas a feature in FFM.</li>\n<li>Hashed all combinations of the document/traffic_source clicked by a\nuser in     page_views. So if a user came to a document from\n'search', it would be treated differently to a user coming to the\nsame document from 'internal'. Any documents occurring less than 80\ntimes in events were dropped.</li>\n<li>Hashed sources (from documents_meta) of all the page view\ndocuments for each user. This gave slight improvement, perhaps due to\nthe fact that it better captured rare documents.</li>\n<li>One hour after clicks – hashed as a new feature the documents clicked\nby a user within one hour of the ad click. This is a form of leak, as\nit is likely after clicking the ad, there are further links in the ad\nto other documents, or users search content related to the    ad.</li>\n<li>Interaction of ad_doc_category and doc_category. Weekdays – worked\nwell in the stack – and hour.</li>\n<li>Log of time difference between display doc creation time and current\ntime; as well as difference between display doc creation time and ad\ndoc creation time Within     page_views, if user viewed ad documents\nof the same publisher, or same source</li>\n<li>Flags, if user viewed ad documents of similar\ncategory, or similar topic.</li>\n<li>Flags, if user viewed this ad in the past, if he\nclicked it in the past; Same for publisher, source, category and\ntopic.</li>\n<li>Flags, if the user viewed this ad in the future, same for ad doc, or if user not\nviewed this ad in the future, but viewed ads of the same campaign.</li>\n</ul>\n\n<p><br><strong>Models used @ L1</strong><br></p>\n\n<ul>\n<li>LibFFM</li>\n<li>Vowpal Wabbit FTRL as well as tingrtu's version</li>\n<li>Liblinear</li>\n<li>XGBoost</li>\n<li>Keras</li>\n<li>Logistic regression</li>\n<li>SVC</li>\n</ul>\n\n<p>As mentioned above, Alexey wrote a custom version of FFM which was much faster and consumed considerably less memory. This will be published soon on github for reuse. </p>\n\n<p><br><strong>CV &amp; Meta Modelling</strong><br>\nWe used a subset of around 6M rows as our validation set, which we sampled according to the structure of the given test test (2 future days, 50% common days / 50% future days rows). Additionally, we used a training subset of around 14M rows in order to be able to test new ideas faster. </p>\n\n<p>Before our team merge with Alexey, stacking around 20 L1-models on this 6M-set has given us a gain of around 0.003 at the public leaderboard. For that, we used a blend of Xgb &amp; Keras models trained on the 6M set at once (=&gt; no common days / future days separation).</p>\n\n<p>During the last week, we improved our single model performance significantly and the gain from stacking decreased to roughly 0.001. On top of the L1-predictions we used normalized time as feature at L2, which yielded a gain of around 0.00020. \nAlexey had its own meta-stack ready by the time we merged, and we blended that into our final submission (both stacks showed low correlation until the very end of the competition, even though Alexey rebuilt his models for our 6M validation set). </p>\n\n<p><br><strong>Final solution</strong><br>\nOur final submission is a geometric mean blend of Alexey’s meta stack, a bagged XGBoost and a bagged Keras model. Weights have been chosen by gut feeling and LB scores (we had no out-of-fold predictions for L2 to optimize the weights on).</p>\n\n<p><br><strong>Best single model</strong><br>\nOur best single model comes from Alexey's custom FFM implementation (2 times bagged). It scores 0.70017 at public LB and 0.70053 at private LB.</p>",
  "messages": [
    {
      "id": "157494",
      "postDate": "01/21/2017 15:53:22",
      "content": "<p>Hi fellow kagglers,</p>\n\n<p>Before starting ... a big thanks to Outbrain and Kaggle for hosting this competition.  Also congratulations to code monkeys for a well deserved win, to Three Data Points and last but not least to my teammates: This has been a really fun competition, the teamwork was just great and there was plenty of things to learn.</p>\n\n<p>Our solution involved a lot of feature engineering, model diversity, stacking and a custum implementation of FFM by Alexey, which allowed very fast iterations and yielded a very strong single model at the end.</p>\n\n<p><br><strong>Our most important features</strong><br></p>\n\n<ul>\n<li>rcarson's leak feature. This was very strong. It was slightly\nimproved by    bucketing it into rows where the page_view doc was\nbefore, 1 hour after, 1 day after and &gt; 1 day after the display\ntimestamp.</li>\n<li>Competing ads – on each row, we hashed each individual competing ad\nas a feature in FFM.</li>\n<li>Hashed all combinations of the document/traffic_source clicked by a\nuser in     page_views. So if a user came to a document from\n'search', it would be treated differently to a user coming to the\nsame document from 'internal'. Any documents occurring less than 80\ntimes in events were dropped.</li>\n<li>Hashed sources (from documents_meta) of all the page view\ndocuments for each user. This gave slight improvement, perhaps due to\nthe fact that it better captured rare documents.</li>\n<li>One hour after clicks – hashed as a new feature the documents clicked\nby a user within one hour of the ad click. This is a form of leak, as\nit is likely after clicking the ad, there are further links in the ad\nto other documents, or users search content related to the    ad.</li>\n<li>Interaction of ad_doc_category and doc_category. Weekdays – worked\nwell in the stack – and hour.</li>\n<li>Log of time difference between display doc creation time and current\ntime; as well as difference between display doc creation time and ad\ndoc creation time Within     page_views, if user viewed ad documents\nof the same publisher, or same source</li>\n<li>Flags, if user viewed ad documents of similar\ncategory, or similar topic.</li>\n<li>Flags, if user viewed this ad in the past, if he\nclicked it in the past; Same for publisher, source, category and\ntopic.</li>\n<li>Flags, if the user viewed this ad in the future, same for ad doc, or if user not\nviewed this ad in the future, but viewed ads of the same campaign.</li>\n</ul>\n\n<p><br><strong>Models used @ L1</strong><br></p>\n\n<ul>\n<li>LibFFM</li>\n<li>Vowpal Wabbit FTRL as well as tingrtu's version</li>\n<li>Liblinear</li>\n<li>XGBoost</li>\n<li>Keras</li>\n<li>Logistic regression</li>\n<li>SVC</li>\n</ul>\n\n<p>As mentioned above, Alexey wrote a custom version of FFM which was much faster and consumed considerably less memory. This will be published soon on github for reuse. </p>\n\n<p><br><strong>CV &amp; Meta Modelling</strong><br>\nWe used a subset of around 6M rows as our validation set, which we sampled according to the structure of the given test test (2 future days, 50% common days / 50% future days rows). Additionally, we used a training subset of around 14M rows in order to be able to test new ideas faster. </p>\n\n<p>Before our team merge with Alexey, stacking around 20 L1-models on this 6M-set has given us a gain of around 0.003 at the public leaderboard. For that, we used a blend of Xgb &amp; Keras models trained on the 6M set at once (=&gt; no common days / future days separation).</p>\n\n<p>During the last week, we improved our single model performance significantly and the gain from stacking decreased to roughly 0.001. On top of the L1-predictions we used normalized time as feature at L2, which yielded a gain of around 0.00020. \nAlexey had its own meta-stack ready by the time we merged, and we blended that into our final submission (both stacks showed low correlation until the very end of the competition, even though Alexey rebuilt his models for our 6M validation set). </p>\n\n<p><br><strong>Final solution</strong><br>\nOur final submission is a geometric mean blend of Alexey’s meta stack, a bagged XGBoost and a bagged Keras model. Weights have been chosen by gut feeling and LB scores (we had no out-of-fold predictions for L2 to optimize the weights on).</p>\n\n<p><br><strong>Best single model</strong><br>\nOur best single model comes from Alexey's custom FFM implementation (2 times bagged). It scores 0.70017 at public LB and 0.70053 at private LB.</p>",
      "rawMarkdown": "Hi fellow kagglers,\r\n\r\nBefore starting ... a big thanks to Outbrain and Kaggle for hosting this competition.  Also congratulations to code monkeys for a well deserved win, to Three Data Points and last but not least to my teammates: This has been a really fun competition, the teamwork was just great and there was plenty of things to learn.\r\n\r\nOur solution involved a lot of feature engineering, model diversity, stacking and a custum implementation of FFM by Alexey, which allowed very fast iterations and yielded a very strong single model at the end.\r\n\r\n<br>**Our most important features**<br>\r\n\r\n - rcarson's leak feature. This was very strong. It was slightly\r\n   improved by    bucketing it into rows where the page_view doc was\r\n   before, 1 hour after, 1 day after and > 1 day after the display\r\n   timestamp.\r\n - Competing ads – on each row, we hashed each individual competing ad\r\n   as a feature in FFM.\r\n - Hashed all combinations of the document/traffic_source clicked by a\r\n   user in     page_views. So if a user came to a document from\r\n   'search', it would be treated differently to a user coming to the\r\n   same document from 'internal'. Any documents occurring less than 80\r\n   times in events were dropped.\r\n - Hashed sources (from documents_meta) of all the page view\r\n   documents for each user. This gave slight improvement, perhaps due to\r\n   the fact that it better captured rare documents.\r\n - One hour after clicks – hashed as a new feature the documents clicked\r\n   by a user within one hour of the ad click. This is a form of leak, as\r\n   it is likely after clicking the ad, there are further links in the ad\r\n   to other documents, or users search content related to the    ad.\r\n - Interaction of ad_doc_category and doc_category. Weekdays – worked\r\n   well in the stack – and hour.\r\n - Log of time difference between display doc creation time and current\r\n   time; as well as difference between display doc creation time and ad\r\n   doc creation time Within     page_views, if user viewed ad documents\r\n   of the same publisher, or same source\r\n - Flags, if user viewed ad documents of similar\r\n   category, or similar topic.\r\n - Flags, if user viewed this ad in the past, if he\r\n   clicked it in the past; Same for publisher, source, category and\r\n   topic.\r\n - Flags, if the user viewed this ad in the future, same for ad doc, or if user not\r\n   viewed this ad in the future, but viewed ads of the same campaign.\r\n \r\n<br>**Models used @ L1**<br>\r\n\r\n - LibFFM\r\n - Vowpal Wabbit FTRL as well as tingrtu's version\r\n - Liblinear\r\n - XGBoost\r\n - Keras\r\n - Logistic regression\r\n - SVC\r\n\r\nAs mentioned above, Alexey wrote a custom version of FFM which was much faster and consumed considerably less memory. This will be published soon on github for reuse. \r\n\r\n<br>**CV & Meta Modelling**<br>\r\nWe used a subset of around 6M rows as our validation set, which we sampled according to the structure of the given test test (2 future days, 50% common days / 50% future days rows). Additionally, we used a training subset of around 14M rows in order to be able to test new ideas faster. \r\n\r\nBefore our team merge with Alexey, stacking around 20 L1-models on this 6M-set has given us a gain of around 0.003 at the public leaderboard. For that, we used a blend of Xgb & Keras models trained on the 6M set at once (=> no common days / future days separation).\r\n\r\n During the last week, we improved our single model performance significantly and the gain from stacking decreased to roughly 0.001. On top of the L1-predictions we used normalized time as feature at L2, which yielded a gain of around 0.00020. \r\nAlexey had its own meta-stack ready by the time we merged, and we blended that into our final submission (both stacks showed low correlation until the very end of the competition, even though Alexey rebuilt his models for our 6M validation set). \r\n\r\n\r\n<br>**Final solution**<br>\r\nOur final submission is a geometric mean blend of Alexey’s meta stack, a bagged XGBoost and a bagged Keras model. Weights have been chosen by gut feeling and LB scores (we had no out-of-fold predictions for L2 to optimize the weights on).\r\n\r\n<br>**Best single model**<br>\r\nOur best single model comes from Alexey's custom FFM implementation (2 times bagged). It scores 0.70017 at public LB and 0.70053 at private LB.",
      "votes": null
    },
    {
      "id": "157551",
      "postDate": "01/22/2017 03:38:10",
      "content": "<p>congratulation and thanks for sharing, would you release your solution code ?</p>",
      "rawMarkdown": "congratulation and thanks for sharing, would you release your solution code ?",
      "votes": null
    },
    {
      "id": "157564",
      "postDate": "01/22/2017 06:20:37",
      "content": "<p>this is awesome, and well deserved</p>",
      "rawMarkdown": "this is awesome, and well deserved",
      "votes": null
    },
    {
      "id": "157580",
      "postDate": "01/22/2017 09:06:49",
      "content": "<p>I've released my part of source code in the current state - <a href=\"https://github.com/alno/kaggle-outbrain-click-prediction\">https://github.com/alno/kaggle-outbrain-click-prediction</a></p>\n\n<p>In particular, it contains code for the best ffm model (it's <strong>ffm2-f4b</strong>, so you may look at <em>export-bin-data-f4.cpp</em> for the features and <em>ffm.cpp</em> and <em>ffm-model.cpp</em> for the ffm implementation).</p>\n\n<p>But as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.</p>",
      "rawMarkdown": "I've released my part of source code in the current state - https://github.com/alno/kaggle-outbrain-click-prediction\r\n\r\nIn particular, it contains code for the best ffm model (it's **ffm2-f4b**, so you may look at *export-bin-data-f4.cpp* for the features and *ffm.cpp* and *ffm-model.cpp* for the ffm implementation).\r\n\r\nBut as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.",
      "votes": null
    },
    {
      "id": "157672",
      "postDate": "01/23/2017 02:18:12",
      "content": "<p>[quote=Alexey Noskov;157580]</p>\n\n<p>I've released my part of source code in the current state - <a href=\"https://github.com/alno/kaggle-outbrain-click-prediction\">https://github.com/alno/kaggle-outbrain-click-prediction</a></p>\n\n<p>In particular, it contains code for the best ffm model (it's <strong>ffm2-f4b</strong>, so you may look at <em>export-bin-data-f4.cpp</em> for the features and <em>ffm.cpp</em> and <em>ffm-model.cpp</em> for the ffm implementation).</p>\n\n<p>But as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.</p>\n\n<p>[/quote]\nawesome material for the kaggle community , looking forward to your next version,thanks again</p>",
      "rawMarkdown": "[quote=Alexey Noskov;157580]\r\n\r\nI've released my part of source code in the current state - https://github.com/alno/kaggle-outbrain-click-prediction\r\n\r\nIn particular, it contains code for the best ffm model (it's **ffm2-f4b**, so you may look at *export-bin-data-f4.cpp* for the features and *ffm.cpp* and *ffm-model.cpp* for the ffm implementation).\r\n\r\nBut as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.\r\n\r\n[/quote]\r\nawesome material for the kaggle community , looking forward to your next version,thanks again",
      "votes": null
    },
    {
      "id": "157725",
      "postDate": "01/23/2017 11:35:35",
      "content": "<p>Many thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. </p>\n\n<p>Our FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. </p>\n\n<p>We sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&amp;wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.</p>\n\n<p>Besides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.</p>",
      "rawMarkdown": "Many thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. \r\n\r\nOur FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. \r\n\r\nWe sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.\r\n\r\nBesides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.",
      "votes": null
    },
    {
      "id": "157737",
      "postDate": "01/23/2017 13:27:39",
      "content": "<p>[quote=nomo;157725]</p>\n\n<p>Many thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. </p>\n\n<p>Our FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. </p>\n\n<p>We sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&amp;wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.</p>\n\n<p>Besides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.</p>\n\n<p>[/quote]\nGreat solutions and congratulations, looking forward to your more details or code release of your solution.</p>",
      "rawMarkdown": "[quote=nomo;157725]\r\n\r\nMany thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. \r\n\r\nOur FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. \r\n\r\nWe sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.\r\n\r\nBesides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.\r\n\r\n[/quote]\r\nGreat solutions and congratulations, looking forward to your more details or code release of your solution.",
      "votes": null
    },
    {
      "id": "158774",
      "postDate": "01/30/2017 01:35:40",
      "content": "<p>Thanks for sharing the code @Alexey Noskov ! It is impressive you could get a score higher than 0.70 with a single FFM model! <br>\nDo you know what would be your score with the same datasets if using standard LibFFM? <br>\nWhat were you improvements on FFM implementation in terms of accuracy? <br>\nWhat does it means that your custom FFM implementation was 2 times bagged (according to @Faron)? <br>\nThanks!</p>",
      "rawMarkdown": "Thanks for sharing the code @Alexey Noskov ! It is impressive you could get a score higher than 0.70 with a single FFM model!  \r\nDo you know what would be your score with the same datasets if using standard LibFFM?  \r\nWhat were you improvements on FFM implementation in terms of accuracy?  \r\nWhat does it means that your custom FFM implementation was 2 times bagged (according to @Faron)?   \r\nThanks!",
      "votes": null
    },
    {
      "id": "160610",
      "postDate": "02/08/2017 18:40:42",
      "content": "<p>Hi!\nIn the features section, what does \"flags\" and \"hashed\" means? I am new in data science and try to learn.\nThanks!</p>",
      "rawMarkdown": "Hi!\nIn the features section, what does \"flags\" and \"hashed\" means? I am new in data science and try to learn.\nThanks!",
      "votes": null
    },
    {
      "id": "160626",
      "postDate": "02/08/2017 21:06:06",
      "content": "<p>Hi Yair, flag is simply an indicator if a case is true or not represented by a number. for example if a flag was \"is_red?\"... the value \"red\", gets turned to 1; while \"blue\" or \"green\" would be turned to 0. Hashing is a means of turning a string to a unique integer value; so it can be used in a ML model. Something like this : <a href=\"http://pythoncentral.io/hashing-strings-with-python/\">http://pythoncentral.io/hashing-strings-with-python/</a> </p>",
      "rawMarkdown": "Hi Yair, flag is simply an indicator if a case is true or not represented by a number. for example if a flag was \"is_red?\"... the value \"red\", gets turned to 1; while \"blue\" or \"green\" would be turned to 0. Hashing is a means of turning a string to a unique integer value; so it can be used in a ML model. Something like this : http://pythoncentral.io/hashing-strings-with-python/",
      "votes": null
    },
    {
      "id": "160628",
      "postDate": "02/08/2017 21:11:39",
      "content": "<p>Hi Gabriel, to my knowledge we did not test the exact same features on the custom FFM (batch-learn) and LibFFM... and alexey may correct me, but I believe the same FFM algorithm is used in both. The real power of batch-learn FFM is the speed and the very low memory it uses. It works off C++ binaries as opposed to LibFFM files, so it loads data very fast and the memory is kept low. See <a href=\"https://github.com/alno/batch-learn\">https://github.com/alno/batch-learn</a> <br>\n2x bagging means ... we ran the model twice using the same features and average the results to stabilise them. </p>",
      "rawMarkdown": "Hi Gabriel, to my knowledge we did not test the exact same features on the custom FFM (batch-learn) and LibFFM... and alexey may correct me, but I believe the same FFM algorithm is used in both. The real power of batch-learn FFM is the speed and the very low memory it uses. It works off C++ binaries as opposed to LibFFM files, so it loads data very fast and the memory is kept low. See https://github.com/alno/batch-learn   \n2x bagging means ... we ran the model twice using the same features and average the results to stabilise them.",
      "votes": null
    },
    {
      "id": "161800",
      "postDate": "02/15/2017 21:35:24",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 157551,
      "author_name": "josephchan",
      "author_url": "",
      "post_date": "01/22/2017 03:38:10",
      "content": "<p>congratulation and thanks for sharing, would you release your solution code ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157564,
      "author_name": "cherednychenko",
      "author_url": "",
      "post_date": "01/22/2017 06:20:37",
      "content": "<p>this is awesome, and well deserved</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157580,
      "author_name": "alexeynoskov",
      "author_url": "",
      "post_date": "01/22/2017 09:06:49",
      "content": "<p>I've released my part of source code in the current state - <a href=\"https://github.com/alno/kaggle-outbrain-click-prediction\">https://github.com/alno/kaggle-outbrain-click-prediction</a></p>\n\n<p>In particular, it contains code for the best ffm model (it's <strong>ffm2-f4b</strong>, so you may look at <em>export-bin-data-f4.cpp</em> for the features and <em>ffm.cpp</em> and <em>ffm-model.cpp</em> for the ffm implementation).</p>\n\n<p>But as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 157672,
          "author_name": "josephchan",
          "author_url": "",
          "post_date": "01/23/2017 02:18:12",
          "content": "<p>[quote=Alexey Noskov;157580]</p>\n\n<p>I've released my part of source code in the current state - <a href=\"https://github.com/alno/kaggle-outbrain-click-prediction\">https://github.com/alno/kaggle-outbrain-click-prediction</a></p>\n\n<p>In particular, it contains code for the best ffm model (it's <strong>ffm2-f4b</strong>, so you may look at <em>export-bin-data-f4.cpp</em> for the features and <em>ffm.cpp</em> and <em>ffm-model.cpp</em> for the ffm implementation).</p>\n\n<p>But as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.</p>\n\n<p>[/quote]\nawesome material for the kaggle community , looking forward to your next version,thanks again</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157725,
      "author_name": "ecow1338",
      "author_url": "",
      "post_date": "01/23/2017 11:35:35",
      "content": "<p>Many thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. </p>\n\n<p>Our FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. </p>\n\n<p>We sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&amp;wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.</p>\n\n<p>Besides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 157737,
          "author_name": "josephchan",
          "author_url": "",
          "post_date": "01/23/2017 13:27:39",
          "content": "<p>[quote=nomo;157725]</p>\n\n<p>Many thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. </p>\n\n<p>Our FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. </p>\n\n<p>We sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&amp;wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.</p>\n\n<p>Besides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.</p>\n\n<p>[/quote]\nGreat solutions and congratulations, looking forward to your more details or code release of your solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 158774,
      "author_name": "gspmoreira",
      "author_url": "",
      "post_date": "01/30/2017 01:35:40",
      "content": "<p>Thanks for sharing the code @Alexey Noskov ! It is impressive you could get a score higher than 0.70 with a single FFM model! <br>\nDo you know what would be your score with the same datasets if using standard LibFFM? <br>\nWhat were you improvements on FFM implementation in terms of accuracy? <br>\nWhat does it means that your custom FFM implementation was 2 times bagged (according to @Faron)? <br>\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 160628,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "02/08/2017 21:11:39",
          "content": "<p>Hi Gabriel, to my knowledge we did not test the exact same features on the custom FFM (batch-learn) and LibFFM... and alexey may correct me, but I believe the same FFM algorithm is used in both. The real power of batch-learn FFM is the speed and the very low memory it uses. It works off C++ binaries as opposed to LibFFM files, so it loads data very fast and the memory is kept low. See <a href=\"https://github.com/alno/batch-learn\">https://github.com/alno/batch-learn</a> <br>\n2x bagging means ... we ran the model twice using the same features and average the results to stabilise them. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 160610,
      "author_name": "yr4000",
      "author_url": "",
      "post_date": "02/08/2017 18:40:42",
      "content": "<p>Hi!\nIn the features section, what does \"flags\" and \"hashed\" means? I am new in data science and try to learn.\nThanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 160626,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "02/08/2017 21:06:06",
          "content": "<p>Hi Yair, flag is simply an indicator if a case is true or not represented by a number. for example if a flag was \"is_red?\"... the value \"red\", gets turned to 1; while \"blue\" or \"green\" would be turned to 0. Hashing is a means of turning a string to a unique integer value; so it can be used in a ML model. Something like this : <a href=\"http://pythoncentral.io/hashing-strings-with-python/\">http://pythoncentral.io/hashing-strings-with-python/</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 161800,
          "author_name": "yr4000",
          "author_url": "",
          "post_date": "02/15/2017 21:35:24",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "157494": "Hi fellow kagglers,\r\n\r\nBefore starting ... a big thanks to Outbrain and Kaggle for hosting this competition.  Also congratulations to code monkeys for a well deserved win, to Three Data Points and last but not least to my teammates: This has been a really fun competition, the teamwork was just great and there was plenty of things to learn.\r\n\r\nOur solution involved a lot of feature engineering, model diversity, stacking and a custum implementation of FFM by Alexey, which allowed very fast iterations and yielded a very strong single model at the end.\r\n\r\n<br>**Our most important features**<br>\r\n\r\n - rcarson's leak feature. This was very strong. It was slightly\r\n   improved by    bucketing it into rows where the page_view doc was\r\n   before, 1 hour after, 1 day after and > 1 day after the display\r\n   timestamp.\r\n - Competing ads – on each row, we hashed each individual competing ad\r\n   as a feature in FFM.\r\n - Hashed all combinations of the document/traffic_source clicked by a\r\n   user in     page_views. So if a user came to a document from\r\n   'search', it would be treated differently to a user coming to the\r\n   same document from 'internal'. Any documents occurring less than 80\r\n   times in events were dropped.\r\n - Hashed sources (from documents_meta) of all the page view\r\n   documents for each user. This gave slight improvement, perhaps due to\r\n   the fact that it better captured rare documents.\r\n - One hour after clicks – hashed as a new feature the documents clicked\r\n   by a user within one hour of the ad click. This is a form of leak, as\r\n   it is likely after clicking the ad, there are further links in the ad\r\n   to other documents, or users search content related to the    ad.\r\n - Interaction of ad_doc_category and doc_category. Weekdays – worked\r\n   well in the stack – and hour.\r\n - Log of time difference between display doc creation time and current\r\n   time; as well as difference between display doc creation time and ad\r\n   doc creation time Within     page_views, if user viewed ad documents\r\n   of the same publisher, or same source\r\n - Flags, if user viewed ad documents of similar\r\n   category, or similar topic.\r\n - Flags, if user viewed this ad in the past, if he\r\n   clicked it in the past; Same for publisher, source, category and\r\n   topic.\r\n - Flags, if the user viewed this ad in the future, same for ad doc, or if user not\r\n   viewed this ad in the future, but viewed ads of the same campaign.\r\n \r\n<br>**Models used @ L1**<br>\r\n\r\n - LibFFM\r\n - Vowpal Wabbit FTRL as well as tingrtu's version\r\n - Liblinear\r\n - XGBoost\r\n - Keras\r\n - Logistic regression\r\n - SVC\r\n\r\nAs mentioned above, Alexey wrote a custom version of FFM which was much faster and consumed considerably less memory. This will be published soon on github for reuse. \r\n\r\n<br>**CV & Meta Modelling**<br>\r\nWe used a subset of around 6M rows as our validation set, which we sampled according to the structure of the given test test (2 future days, 50% common days / 50% future days rows). Additionally, we used a training subset of around 14M rows in order to be able to test new ideas faster. \r\n\r\nBefore our team merge with Alexey, stacking around 20 L1-models on this 6M-set has given us a gain of around 0.003 at the public leaderboard. For that, we used a blend of Xgb & Keras models trained on the 6M set at once (=> no common days / future days separation).\r\n\r\n During the last week, we improved our single model performance significantly and the gain from stacking decreased to roughly 0.001. On top of the L1-predictions we used normalized time as feature at L2, which yielded a gain of around 0.00020. \r\nAlexey had its own meta-stack ready by the time we merged, and we blended that into our final submission (both stacks showed low correlation until the very end of the competition, even though Alexey rebuilt his models for our 6M validation set). \r\n\r\n\r\n<br>**Final solution**<br>\r\nOur final submission is a geometric mean blend of Alexey’s meta stack, a bagged XGBoost and a bagged Keras model. Weights have been chosen by gut feeling and LB scores (we had no out-of-fold predictions for L2 to optimize the weights on).\r\n\r\n<br>**Best single model**<br>\r\nOur best single model comes from Alexey's custom FFM implementation (2 times bagged). It scores 0.70017 at public LB and 0.70053 at private LB.",
    "157551": "congratulation and thanks for sharing, would you release your solution code ?",
    "157564": "this is awesome, and well deserved",
    "157580": "I've released my part of source code in the current state - https://github.com/alno/kaggle-outbrain-click-prediction\r\n\r\nIn particular, it contains code for the best ffm model (it's **ffm2-f4b**, so you may look at *export-bin-data-f4.cpp* for the features and *ffm.cpp* and *ffm-model.cpp* for the ffm implementation).\r\n\r\nBut as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.",
    "157672": "[quote=Alexey Noskov;157580]\r\n\r\nI've released my part of source code in the current state - https://github.com/alno/kaggle-outbrain-click-prediction\r\n\r\nIn particular, it contains code for the best ffm model (it's **ffm2-f4b**, so you may look at *export-bin-data-f4.cpp* for the features and *ffm.cpp* and *ffm-model.cpp* for the ffm implementation).\r\n\r\nBut as it uses some features which i didn't generate by myself, but received from my teammates, it's hard to run it now. Some time later we'll merge required code for feature generation and then it will be possible to run it.\r\n\r\n[/quote]\r\nawesome material for the kaggle community , looking forward to your next version,thanks again",
    "157725": "Many thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. \r\n\r\nOur FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. \r\n\r\nWe sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.\r\n\r\nBesides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.",
    "157737": "[quote=nomo;157725]\r\n\r\nMany thanks to brain-afk's kindly share, your algorithm and solution is really awesome. Actually, we are very amazed by your dramatically improvement of last period. We are just a little lucky in the end. \r\n\r\nOur FFM model main features are quite similar. Some of your fine-grained feature engineering process should be better. My teammate hugepanda implements our own version FFM and LR based on libffm with lots of optimization. We use around 1/10 data first for the fast iteration and build up several models of different features to produce diversity, i.e., generalized ffm models, memorilized manually crossed LR models, statistics based boosting tree models. \r\n\r\nWe sampled validation data and stacking layer 2(L2) training data according to test distribution. Validation is about 15%, L2 is about 35%, and the remainining 50% is for L1 single model. We first train single models on L1 data and predict on L2 data, the L2 model is trained using L1 prediction. We then use L1+L2 data to retrain new L1 model for the L2 prediction on validation data(this result is also used in another stacking training), use full data to retrain new L1 model for L2 prediction on test data, which could improve our result a lot. Our teammates Bowen-Yuan, Eureka and Daniel also bring powerful ffms with interesting features, together with our single models and L2 model result to do another very useful and interesting validation data based two-stage stacking(nn,boosting tree,linear model), and then bagged again with aggregated several L2 models to generate the final ranking result. Especially last several days we collaberate very well and inspire each other a lot. We have brain-storming and new ideas everyday. This is definitely a cool experience. Also we have many unsuccessful experiments, like matrix decomposition, ID feature multi-layer perceptron(e.g., google apps deep&wide network, youtube shared weight network structure), any successful try or comment of them is very welcome.\r\n\r\nBesides, we're quite inspired by so many brilliant algorithm of kagglers' share these days. Congratuations to brain-afk and Three Data Points team. The leak and scripts kindly shared by Three Data Points team is quite useful for our improvement. Thanks libffm, liblinear, tensorflow, keras, lightgbm, xgboost for the great tool. Last but not least, I'd like to thank my teammates hugepanda, Bowen-Yuan, Eureka and Daniel, and I have learnt a lot from them.\r\n\r\n[/quote]\r\nGreat solutions and congratulations, looking forward to your more details or code release of your solution.",
    "158774": "Thanks for sharing the code @Alexey Noskov ! It is impressive you could get a score higher than 0.70 with a single FFM model!  \r\nDo you know what would be your score with the same datasets if using standard LibFFM?  \r\nWhat were you improvements on FFM implementation in terms of accuracy?  \r\nWhat does it means that your custom FFM implementation was 2 times bagged (according to @Faron)?   \r\nThanks!",
    "160610": "Hi!\nIn the features section, what does \"flags\" and \"hashed\" means? I am new in data science and try to learn.\nThanks!",
    "160626": "Hi Yair, flag is simply an indicator if a case is true or not represented by a number. for example if a flag was \"is_red?\"... the value \"red\", gets turned to 1; while \"blue\" or \"green\" would be turned to 0. Hashing is a means of turning a string to a unique integer value; so it can be used in a ML model. Something like this : http://pythoncentral.io/hashing-strings-with-python/",
    "160628": "Hi Gabriel, to my knowledge we did not test the exact same features on the custom FFM (batch-learn) and LibFFM... and alexey may correct me, but I believe the same FFM algorithm is used in both. The real power of batch-learn FFM is the speed and the very low memory it uses. It works off C++ binaries as opposed to LibFFM files, so it loads data very fast and the memory is kept low. See https://github.com/alno/batch-learn   \n2x bagging means ... we ran the model twice using the same features and average the results to stabilise them.",
    "161800": "Thank you!"
  },
  "source": "meta"
}