{
  "id": 27923,
  "title": "3rd place solution",
  "url": "/competitions/outbrain-click-prediction/discussion/27923",
  "author_name": "tito",
  "post_date": "2017-01-19T14:18:11.927000",
  "votes": 21,
  "comment_count": 17,
  "views": 766,
  "content": "<p>Congratulations to code monkey, brain-afk and everyone!\nAnd I’d like to thank my teammates carl and Little Boat. I could learn lots of things from them.</p>\n\n<p>As <a href=\"https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27897/congrats-and-solution-sharing?forumMessageId=157062#post157062\">Little Boat explained</a>, our model is ffm models + xgboost models as first layer, and xgboost as second layer.</p>\n\n<p>I would like to explain our first layer models here.</p>\n\n<h3><strong>Feature extraction</strong></h3>\n\n<p>Following is the list of unobvious features we used.\nWe made lots of other features which turned out to be useless.\nBut analyzing large scale user behavior was very interesting to me.</p>\n\n<ul>\n<li>page view counts of each user.</li>\n<li>page view counts of ad landing page.</li>\n<li>impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.</li>\n<li>landing page confidence vector, where confidence vector is defined as a vector which element is composed of confidence level from documents_*.csv. This feature is used only for ffm as numeric data.</li>\n<li>user confidence vector - this is the average of document confidence vector which is viewed by each user. This feature is used only for ffm as numeric data.</li>\n<li>inner dot product of document confidence vectors of ad impression page and ad landing page .</li>\n<li>inner dot product of user confidence vector and ad landing page document confidence vector</li>\n<li>XGB leaf for ffm feature.</li>\n<li>immediate document viewed after click event.</li>\n</ul>\n\n<h3><strong>Modeling and Training</strong></h3>\n\n<strong>FFM</strong>\n\n<p>In addition to <a href=\"https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27892/introducing-light-ffm-and-stack-nn\">carl’s light-ffm</a>, we used another customized libffm which has pair-wise rank and can take sample weight.\nBest public LB score is 0.6974.</p>\n\n<strong>XGB</strong>\n\n<p>We used xgboost with rank:pairwise objective.\nBest public LB score is 0.6885.</p>\n\n<strong>Leak Row Excluding</strong>\n\n<p>We trained ffm models without leak row by removing them from training data first, then blend/merge leak information later.\nThis improved single model score.</p>\n\n<strong>Present/Future split</strong>\n\n<p>We split test data into ‘present’ data (in time samples) and ‘future’ data (out time samples), then built models for each data.\nHere, Features and hyper parameters for each model are optimized independently.\nFor example we used timestamp feature for ‘present’ model, but did not used it for ‘future’ model. And we used sample weight of training data for ‘future’ model to add higher weight to last day than first day.</p>\n\n<p>Please feel free to ask any questions.</p>",
  "messages": [
    {
      "id": 157166,
      "postDate": "2017-01-19T14:18:11.927Z",
      "content": "<p>Congratulations to code monkey, brain-afk and everyone!\nAnd I’d like to thank my teammates carl and Little Boat. I could learn lots of things from them.</p>\n\n<p>As <a href=\"https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27897/congrats-and-solution-sharing?forumMessageId=157062#post157062\">Little Boat explained</a>, our model is ffm models + xgboost models as first layer, and xgboost as second layer.</p>\n\n<p>I would like to explain our first layer models here.</p>\n\n<h3><strong>Feature extraction</strong></h3>\n\n<p>Following is the list of unobvious features we used.\nWe made lots of other features which turned out to be useless.\nBut analyzing large scale user behavior was very interesting to me.</p>\n\n<ul>\n<li>page view counts of each user.</li>\n<li>page view counts of ad landing page.</li>\n<li>impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.</li>\n<li>landing page confidence vector, where confidence vector is defined as a vector which element is composed of confidence level from documents_*.csv. This feature is used only for ffm as numeric data.</li>\n<li>user confidence vector - this is the average of document confidence vector which is viewed by each user. This feature is used only for ffm as numeric data.</li>\n<li>inner dot product of document confidence vectors of ad impression page and ad landing page .</li>\n<li>inner dot product of user confidence vector and ad landing page document confidence vector</li>\n<li>XGB leaf for ffm feature.</li>\n<li>immediate document viewed after click event.</li>\n</ul>\n\n<h3><strong>Modeling and Training</strong></h3>\n\n<strong>FFM</strong>\n\n<p>In addition to <a href=\"https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27892/introducing-light-ffm-and-stack-nn\">carl’s light-ffm</a>, we used another customized libffm which has pair-wise rank and can take sample weight.\nBest public LB score is 0.6974.</p>\n\n<strong>XGB</strong>\n\n<p>We used xgboost with rank:pairwise objective.\nBest public LB score is 0.6885.</p>\n\n<strong>Leak Row Excluding</strong>\n\n<p>We trained ffm models without leak row by removing them from training data first, then blend/merge leak information later.\nThis improved single model score.</p>\n\n<strong>Present/Future split</strong>\n\n<p>We split test data into ‘present’ data (in time samples) and ‘future’ data (out time samples), then built models for each data.\nHere, Features and hyper parameters for each model are optimized independently.\nFor example we used timestamp feature for ‘present’ model, but did not used it for ‘future’ model. And we used sample weight of training data for ‘future’ model to add higher weight to last day than first day.</p>\n\n<p>Please feel free to ask any questions.</p>",
      "rawMarkdown": "Congratulations to code monkey, brain-afk and everyone!\r\nAnd I’d like to thank my teammates carl and Little Boat. I could learn lots of things from them.\r\n\r\nAs [Little Boat explained](https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27897/congrats-and-solution-sharing?forumMessageId=157062#post157062), our model is ffm models + xgboost models as first layer, and xgboost as second layer.\r\n\r\nI would like to explain our first layer models here.\r\n\r\n###**Feature extraction**\r\n\r\nFollowing is the list of unobvious features we used.\r\nWe made lots of other features which turned out to be useless.\r\nBut analyzing large scale user behavior was very interesting to me.\r\n\r\n - page view counts of each user.\r\n - page view counts of ad landing page.\r\n - impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.\r\n - landing page confidence vector, where confidence vector is defined as a vector which element is composed of confidence level from documents_*.csv. This feature is used only for ffm as numeric data.\r\n - user confidence vector - this is the average of document confidence vector which is viewed by each user. This feature is used only for ffm as numeric data.\r\n - inner dot product of document confidence vectors of ad impression page and ad landing page .\r\n - inner dot product of user confidence vector and ad landing page document confidence vector\r\n - XGB leaf for ffm feature.\r\n - immediate document viewed after click event.\r\n\r\n###**Modeling and Training**\r\n\r\n####**FFM**\r\nIn addition to [carl’s light-ffm](https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27892/introducing-light-ffm-and-stack-nn), we used another customized libffm which has pair-wise rank and can take sample weight.\r\nBest public LB score is 0.6974.\r\n\r\n####**XGB**\r\n We used xgboost with rank:pairwise objective.\r\nBest public LB score is 0.6885.\r\n\r\n####**Leak Row Excluding**\r\nWe trained ffm models without leak row by removing them from training data first, then blend/merge leak information later.\r\nThis improved single model score.\r\n\r\n####**Present/Future split**\r\nWe split test data into ‘present’ data (in time samples) and ‘future’ data (out time samples), then built models for each data.\r\nHere, Features and hyper parameters for each model are optimized independently.\r\nFor example we used timestamp feature for ‘present’ model, but did not used it for ‘future’ model. And we used sample weight of training data for ‘future’ model to add higher weight to last day than first day.\r\n\r\nPlease feel free to ask any questions.\r\n\r\n",
      "votes": 21
    },
    {
      "id": 158945,
      "postDate": "2017-01-30T22:13:13.557Z",
      "content": "<p>Great @CuteChibiko ! Thanks for sharing!</p>",
      "rawMarkdown": "Great @CuteChibiko ! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 157578,
      "postDate": "2017-01-22T08:57:13.290Z",
      "content": "<p>Thank you Yuriy Makarov,</p>\n\n<p>Yes, we tried CTR with some impression threshold. If impression threshold was low, CTR caused overfit. And if impression threshold was high, CTR did not contribute to score. So We did not use CTR in the end.</p>\n\n<p>\"immediate document viewed after click event.\" means the first document that was seen by a user after the user clicked an ad.  </p>",
      "rawMarkdown": "Thank you Yuriy Makarov,\r\n\r\nYes, we tried CTR with some impression threshold. If impression threshold was low, CTR caused overfit. And if impression threshold was high, CTR did not contribute to score. So We did not use CTR in the end.\r\n\r\n\"immediate document viewed after click event.\" means the first document that was seen by a user after the user clicked an ad.  ",
      "votes": 1
    },
    {
      "id": 157274,
      "postDate": "2017-01-19T23:48:05.547Z",
      "content": "<p>@Gabriel Moreira,</p>\n\n<p>Your CV score drop, from CV=0.679 to 0.670, is very big, so I think something should be happening.\nI'm not sure why, but maybe you are right. We made XGB models for test and validation respectively, so this may cause the difference.</p>\n\n<p>BTW, our best single score did not improved with GBDT leaves, but it worked for ensemble model.</p>\n\n<p>Our best single CV score without GBDT leaves was 0.69586 and score with GBDT leaves was 0.69588 when we tested GBDT leaves. And weighted average of them scored 0.6971.So even single model score was not improved, GBDT leaves feature was very important.</p>",
      "rawMarkdown": "@Gabriel Moreira,\r\n\r\nYour CV score drop, from CV=0.679 to 0.670, is very big, so I think something should be happening.\r\nI'm not sure why, but maybe you are right. We made XGB models for test and validation respectively, so this may cause the difference.\r\n\r\nBTW, our best single score did not improved with GBDT leaves, but it worked for ensemble model.\r\n\r\nOur best single CV score without GBDT leaves was 0.69586 and score with GBDT leaves was 0.69588 when we tested GBDT leaves. And weighted average of them scored 0.6971.So even single model score was not improved, GBDT leaves feature was very important.",
      "votes": 1
    },
    {
      "id": 157225,
      "postDate": "2017-01-19T18:49:26.890Z",
      "content": "<p>@CuteChibiko I also tried the 3 idiots approach to add GBDT leaves nodes as features to FFM, but did not succeeded improving the score using standard LibFFM, (actually they made my FFM CV score worse). I tried 30 trees with 64 and 128 leaves, using XGBoost and LightGBM.\nDo you think my GBDT model may not have being accurate enought (CV=0.670), while my FFM model was (0.679)?\nOther aspect is that I used a fixed validation set following test set days distribution. So, I trained GBDT on full train set and used the model to predict both the train set and validation set. Maybe that could cause some overfit on FFM model, even with using leaf nodes?</p>",
      "rawMarkdown": "@CuteChibiko I also tried the 3 idiots approach to add GBDT leaves nodes as features to FFM, but did not succeeded improving the score using standard LibFFM, (actually they made my FFM CV score worse). I tried 30 trees with 64 and 128 leaves, using XGBoost and LightGBM.\r\nDo you think my GBDT model may not have being accurate enought (CV=0.670), while my FFM model was (0.679)?\r\nOther aspect is that I used a fixed validation set following test set days distribution. So, I trained GBDT on full train set and used the model to predict both the train set and validation set. Maybe that could cause some overfit on FFM model, even with using leaf nodes?",
      "votes": 1
    },
    {
      "id": 157180,
      "postDate": "2017-01-19T15:18:38.207Z",
      "content": "<p>\"impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.\"</p>\n\n<p>what do you mean by impression counts?</p>",
      "rawMarkdown": "\"impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.\"\r\n\r\nwhat do you mean by impression counts?",
      "votes": 1
    },
    {
      "id": 157171,
      "postDate": "2017-01-19T14:40:51.577Z",
      "content": "<p>I have 2 questions: </p>\n\n<ol>\n<li>Could you explain what XGB leaf feature is?  </li>\n<li>Do you normalize the confidence vectors for documents?</li>\n</ol>\n\n<p>Thanks and congrats!</p>",
      "rawMarkdown": "I have 2 questions: \r\n\r\n 1. Could you explain what XGB leaf feature is?  \r\n 2. Do you normalize the confidence vectors for documents?\r\n\r\nThanks and congrats!",
      "votes": 1
    },
    {
      "id": 159127,
      "postDate": "2017-02-01T00:06:25.283Z",
      "content": "<p>Good job @CuteChibiko - I took a look in your implementation and saw it really deserved a try. I've run today a benchmark using some of my datasets in FFM format comparing standard LibFFM and your fork with custom pairwise rank customization. <br>\nI just run the models with the same FFM hyperparameters (without tunning for MAP ranking) and weight 1 for all samples. The environment was a Google Compute Engine instance with 32 CPUs, 256 GB RAM and FFM data in a SSD disk.  </p>\n\n<p>1) Dataset with categorical fields and some binned numeric fields <br>\nParameters: -l 0.0001 -k 6 -t 20 -r 0.2 -s 28 --auto-stop <br>\nRounds: 20  </p>\n\n<p>LibFFM - Cross-validation MAP             = 0.65911 <br>\nLibFFM - Cross-validation MAP + leak = 0.67795 <br>\nLibFFM - Training time                           = 291m  </p>\n\n<p>LibFFM pairwise - Cross-validation MAP             = 0.66040 <br>\nLibFFM pairwise - Cross-validation MAP + leak = 0.67792 <br>\nLibFFM pairwise - Training time                           = 631m  </p>\n\n<hr>\n\n<p>2) Dataset with only categorical fields\nParameters: -l 0.00002 -k 6 -t 50 -r 0.1 -s 28 --auto-stop</p>\n\n<p>LibFFM - Training rounds = 13  (auto-stop on validation logloss) <br>\nLibFFM - Cross-validation MAP             = 0.66428 <br>\nLibFFM - Cross-validation MAP + leak = 0.67859 <br>\nLibFFM - Training time                           = 27m</p>\n\n<p>LibFFM - Training rounds = 26 (auto-stop on validation MAP) <br>\nLibFFM pairwise - Cross-validation MAP             = 0.66552 <br>\nLibFFM pairwise - Cross-validation MAP + leak = 0.67973 <br>\nLibFFM pairwise - Training time                           = 211m</p>\n\n<hr>\n\n<p>I noticed a diff of around 0.001 in CV MAP by for two datasets. Was this the order of difference you could find in your experiments with your FFM with ranking optimization? <br>\nPs. I also noticed a great increase in training logloss (from 0.41009 to 0.56855), but this is expected as the optimization is now driven by MAP and not logloss), right?  </p>\n\n<p>I am worried only about accuracy at this time, but I could notice that training time of your implementation is very slow compared to standard FFM, and that during most of the processing it is using only 1 CPU. Is that because you had not time to parallelize MAP evaluation?   </p>\n\n<p>Which sample weight strategy have you used? More weight on recent samples / clicked samples / display_ids with more ads?  </p>\n\n<p>Thanks again!  </p>",
      "rawMarkdown": "Good job @CuteChibiko - I took a look in your implementation and saw it really deserved a try. I've run today a benchmark using some of my datasets in FFM format comparing standard LibFFM and your fork with custom pairwise rank customization.    \r\nI just run the models with the same FFM hyperparameters (without tunning for MAP ranking) and weight 1 for all samples. The environment was a Google Compute Engine instance with 32 CPUs, 256 GB RAM and FFM data in a SSD disk.  \r\n\r\n1) Dataset with categorical fields and some binned numeric fields  \r\nParameters: -l 0.0001 -k 6 -t 20 -r 0.2 -s 28 --auto-stop  \r\nRounds: 20  \r\n  \r\nLibFFM - Cross-validation MAP             = 0.65911  \r\nLibFFM - Cross-validation MAP + leak = 0.67795  \r\nLibFFM - Training time                           = 291m  \r\n  \r\nLibFFM pairwise - Cross-validation MAP             = 0.66040  \r\nLibFFM pairwise - Cross-validation MAP + leak = 0.67792  \r\nLibFFM pairwise - Training time                           = 631m  \r\n\r\n----------------------------------------------------\r\n\r\n2) Dataset with only categorical fields\r\nParameters: -l 0.00002 -k 6 -t 50 -r 0.1 -s 28 --auto-stop\r\n\r\nLibFFM - Training rounds = 13  (auto-stop on validation logloss)  \r\nLibFFM - Cross-validation MAP             = 0.66428  \r\nLibFFM - Cross-validation MAP + leak = 0.67859  \r\nLibFFM - Training time                           = 27m\r\n\r\nLibFFM - Training rounds = 26 (auto-stop on validation MAP)  \r\nLibFFM pairwise - Cross-validation MAP             = 0.66552  \r\nLibFFM pairwise - Cross-validation MAP + leak = 0.67973  \r\nLibFFM pairwise - Training time                           = 211m\r\n\r\n-----------------------------------\r\n  \r\nI noticed a diff of around 0.001 in CV MAP by for two datasets. Was this the order of difference you could find in your experiments with your FFM with ranking optimization?   \r\nPs. I also noticed a great increase in training logloss (from 0.41009 to 0.56855), but this is expected as the optimization is now driven by MAP and not logloss), right?  \r\n   \r\nI am worried only about accuracy at this time, but I could notice that training time of your implementation is very slow compared to standard FFM, and that during most of the processing it is using only 1 CPU. Is that because you had not time to parallelize MAP evaluation?   \r\n\r\nWhich sample weight strategy have you used? More weight on recent samples / clicked samples / display_ids with more ads?  \r\n\r\nThanks again!  ",
      "votes": 2
    },
    {
      "id": 158836,
      "postDate": "2017-01-30T11:06:15.573Z",
      "content": "<p>@Gabriel Moreira,</p>\n\n<p>I uploaded this source code. You can find pairwise rank customization <a href=\"https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L443\">around here</a>.</p>\n\n<pre><code>if(param.use_default_train){\n  # default libffm training\n} else {\n  # pairwise rank training\n}\n</code></pre>",
      "rawMarkdown": "@Gabriel Moreira,\r\n\r\nI uploaded this source code. You can find pairwise rank customization [around here](https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L443).\r\n\r\n    if(param.use_default_train){\r\n      # default libffm training\r\n    } else {\r\n      # pairwise rank training\r\n    }",
      "votes": 2
    },
    {
      "id": 157187,
      "postDate": "2017-01-19T15:40:34.513Z",
      "content": "<p>@Sameh Faidi,</p>\n\n<p>I meant impression count as how many times the ad is shown.\nExplanation of impression in the context of online advertising can be found <a href=\"https://en.wikipedia.org/wiki/Impression_(online_media)\">wikipedia</a>. </p>\n\n<p>For example, impression counts for advertiser_id means how many times ads from the advertiser is shown.</p>",
      "rawMarkdown": "@Sameh Faidi,\r\n\r\nI meant impression count as how many times the ad is shown.\r\nExplanation of impression in the context of online advertising can be found [wikipedia](https://en.wikipedia.org/wiki/Impression_(online_media)). \r\n\r\nFor example, impression counts for advertiser_id means how many times ads from the advertiser is shown.\r\n",
      "votes": 2
    },
    {
      "id": 157179,
      "postDate": "2017-01-19T15:10:08.563Z",
      "content": "<p>Thank you, Sangzia,</p>\n\n<ol>\n<li>Page 9 of <a href=\"http://www.csie.ntu.edu.tw/~r01922136/kaggle-2014-criteo.pdf\">this presentation by 3 idiots</a> is very good explanation of XGB leaf feature.  </li>\n<li>Document confidence vectors are not normalized, but user confidence vectors are normalized.</li>\n</ol>",
      "rawMarkdown": "Thank you, Sangzia,\r\n\r\n 1. Page 9 of [this presentation by 3 idiots](http://www.csie.ntu.edu.tw/~r01922136/kaggle-2014-criteo.pdf) is very good explanation of XGB leaf feature.  \r\n 2. Document confidence vectors are not normalized, but user confidence vectors are normalized.",
      "votes": 2
    },
    {
      "id": 159411,
      "postDate": "2017-02-02T05:57:11.737Z",
      "content": "<p>Thanks for your kindly share@CuteChibiko . It is really impressive and interesting. I have several questoins about gbdt leaf node as feature.</p>\n\n<p>1). Does the gbdt feature contain label information? like click/ctr. Or if not contain, like impression/pageview count etc.</p>\n\n<p>2). Does the gbdt training data have overlap with ffm training data, If yes, does it suffer from overfitting at first due to label twice used and do you use some method to avoid overfitting?</p>\n\n<p>3). Your weighted average with gbdt leaf's gain is large. Have you compared with ffm containing the gbdt's count feature in some transformation like log without gbdt leafnode, or use gbdt and ffm pctr prediction score together feeded into stacking model. I'm curious that the gain is from gbdt leaf node transformation method or the count/statistics feature itself.</p>\n\n<p>4). Ususallly gbdt training has several hundred or thousand trees, and the cost is large for ffm. How many trees you have used and do you have some hyper-paramter or feature tuning for the less tree numbers?</p>\n\n<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks for your kindly share@CuteChibiko . It is really impressive and interesting. I have several questoins about gbdt leaf node as feature.\n\n1). Does the gbdt feature contain label information? like click/ctr. Or if not contain, like impression/pageview count etc.\n\n2). Does the gbdt training data have overlap with ffm training data, If yes, does it suffer from overfitting at first due to label twice used and do you use some method to avoid overfitting?\n\n3). Your weighted average with gbdt leaf's gain is large. Have you compared with ffm containing the gbdt's count feature in some transformation like log without gbdt leafnode, or use gbdt and ffm pctr prediction score together feeded into stacking model. I'm curious that the gain is from gbdt leaf node transformation method or the count/statistics feature itself.\n\n4). Ususallly gbdt training has several hundred or thousand trees, and the cost is large for ffm. How many trees you have used and do you have some hyper-paramter or feature tuning for the less tree numbers?\n\n\nThanks a lot.\n",
      "replies": [
        {
          "id": 159463,
          "postDate": "2017-02-02T11:35:06.360Z",
          "content": "<p>@nomo,</p>\n\n<p>1) The gbdt feature dose not contain label information but contains impression/pageview count. This is almost same as what I listed in the first post of this topic.</p>\n\n<p>2) Yes, it dose. The gbdt training data was same as ffm training data. But it did not suffer from overfitting.</p>\n\n<p>3) I'm not sure I could understand your question correctly, but we only tried xgb leafnode information itself as feature and did not compared with any variations.</p>\n\n<p>4) We used ntrees=39 and maxDepth=15. Almost all features and hyper-paramters of this model are the same as values of best xgb single model. But only eta was tuned for this model to reduce ntrees.</p>",
          "rawMarkdown": "@nomo,\n\n1) The gbdt feature dose not contain label information but contains impression/pageview count. This is almost same as what I listed in the first post of this topic.\n\n2) Yes, it dose. The gbdt training data was same as ffm training data. But it did not suffer from overfitting.\n\n3) I'm not sure I could understand your question correctly, but we only tried xgb leafnode information itself as feature and did not compared with any variations.\n\n4) We used ntrees=39 and maxDepth=15. Almost all features and hyper-paramters of this model are the same as values of best xgb single model. But only eta was tuned for this model to reduce ntrees."
        },
        {
          "id": 159490,
          "postDate": "2017-02-02T14:44:01.707Z",
          "content": "<p>@CuteChibiko thank you very much for your detailed explain. It is interesting:-)</p>\n\n<p>I have also tried split data as present/future for xgb stacking training, but the result is neutral to not split one, and am interested in the different and your success trying reason.</p>\n\n<p>1). For a.single model ffm / b.single model xgb / c. stacking model, your present/future split has gains on all of these three models or some of these models?</p>\n\n<p>2). How much gain do your solution has compared with not split modeling? Is it mainly from present part or future part? </p>\n\n<p>3). The future model process is interesting. Does the sample weight use exponent decay, like DecayFactor ^ (LastDay - <strong>SampleDay</strong>) as sample weight?</p>\n\n<p>4). I have also used timestamp feature for split modeling present part, but use same hyper-parameter for the present/future part, and the result is neutral. Do your different parameters for these two part influence a lot? Or some other feature or model method also make a difference, like weighted average of present/future split modeling and not split model.</p>\n\n<p>thanks a lot!</p>",
          "rawMarkdown": "@CuteChibiko thank you very much for your detailed explain. It is interesting:-)\n\nI have also tried split data as present/future for xgb stacking training, but the result is neutral to not split one, and am interested in the different and your success trying reason.\n\n1). For a.single model ffm / b.single model xgb / c. stacking model, your present/future split has gains on all of these three models or some of these models?\n\n2). How much gain do your solution has compared with not split modeling? Is it mainly from present part or future part? \n\n3). The future model process is interesting. Does the sample weight use exponent decay, like DecayFactor ^ (LastDay - **SampleDay**) as sample weight?\n\n4). I have also used timestamp feature for split modeling present part, but use same hyper-parameter for the present/future part, and the result is neutral. Do your different parameters for these two part influence a lot? Or some other feature or model method also make a difference, like weighted average of present/future split modeling and not split model.\n\nthanks a lot!\n\n\n"
        },
        {
          "id": 159552,
          "postDate": "2017-02-02T18:48:33.603Z",
          "content": "<p>@nomo,</p>\n\n<p>1) Splitting data worked for a and b. And It seemed to work for validation data set of stacking model, but LB score decreased. We did not have enough time to find the reason, so we did not use splitting for stacking model.</p>\n\n<p>2) The single model gain is 0.001 for ffm future model with sample weight, 0.002 for ffm present model with timestamp feature and 0.002 for xgb present model with timestamp.</p>\n\n<p>3) Exactly. We used this formula: </p>\n\n<pre><code>sample weight = 1 + 2 * 0.5 ^ ((max_timestamp - timestamp)/86400000.0)\n</code></pre>\n\n<p>4) Timestamp feature for present data made big gain without any special method. We used almost same hyper-parameters for split model. Only tuned hyper-parameter is iterations/ntrees.</p>",
          "rawMarkdown": "@nomo,\n\n1) Splitting data worked for a and b. And It seemed to work for validation data set of stacking model, but LB score decreased. We did not have enough time to find the reason, so we did not use splitting for stacking model.\n\n2) The single model gain is 0.001 for ffm future model with sample weight, 0.002 for ffm present model with timestamp feature and 0.002 for xgb present model with timestamp.\n\n3) Exactly. We used this formula: \n\n    sample weight = 1 + 2 * 0.5 ^ ((max_timestamp - timestamp)/86400000.0)\n\n4) Timestamp feature for present data made big gain without any special method. We used almost same hyper-parameters for split model. Only tuned hyper-parameter is iterations/ntrees.\n"
        }
      ]
    },
    {
      "id": 159137,
      "postDate": "2017-02-01T01:45:28.363Z",
      "content": "<p>@Gabriel Moreira, thank you for your feedback.</p>\n\n<p>The difference was 0.0015 ~ 0.002 for our models.\nRelatively bigger ETA for LibFFM pairwise made better results.</p>\n\n<p>You are right. It is using only 1 CPU to calculate validation score.\nI don't remember why I commented out <a href=\"https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L500\">here</a> and did not use omp. :(</p>\n\n<p>We added higher weights to recent samples for ‘future’ model. This improved 'future' model score about 0.001.</p>",
      "rawMarkdown": "@Gabriel Moreira, thank you for your feedback.\r\n\r\nThe difference was 0.0015 ~ 0.002 for our models.\r\nRelatively bigger ETA for LibFFM pairwise made better results.\r\n\r\nYou are right. It is using only 1 CPU to calculate validation score.\r\nI don't remember why I commented out [here](https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L500) and did not use omp. :(\r\n\r\nWe added higher weights to recent samples for ‘future’ model. This improved 'future' model score about 0.001."
    },
    {
      "id": 158778,
      "postDate": "2017-01-30T01:48:11.027Z",
      "content": "<p>Thanks for sharing @CuteChibiko ! You've told you used a customized libffm supporting pairwise rank.\nI am really interested in how the FFM optimization works when using pairwise ranking for libffm.... could you share something about that?</p>",
      "rawMarkdown": "Thanks for sharing @CuteChibiko ! You've told you used a customized libffm supporting pairwise rank.\r\nI am really interested in how the FFM optimization works when using pairwise ranking for libffm.... could you share something about that?"
    },
    {
      "id": 157575,
      "postDate": "2017-01-22T08:01:42.010Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 158945,
      "author_name": "Gabriel Moreira",
      "author_url": "",
      "post_date": "2017-01-30T22:13:13.557000",
      "content": "<p>Great @CuteChibiko ! Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 157578,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2017-01-22T08:57:13.290000",
      "content": "<p>Thank you Yuriy Makarov,</p>\n\n<p>Yes, we tried CTR with some impression threshold. If impression threshold was low, CTR caused overfit. And if impression threshold was high, CTR did not contribute to score. So We did not use CTR in the end.</p>\n\n<p>\"immediate document viewed after click event.\" means the first document that was seen by a user after the user clicked an ad.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 157274,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2017-01-19T23:48:05.547000",
      "content": "<p>@Gabriel Moreira,</p>\n\n<p>Your CV score drop, from CV=0.679 to 0.670, is very big, so I think something should be happening.\nI'm not sure why, but maybe you are right. We made XGB models for test and validation respectively, so this may cause the difference.</p>\n\n<p>BTW, our best single score did not improved with GBDT leaves, but it worked for ensemble model.</p>\n\n<p>Our best single CV score without GBDT leaves was 0.69586 and score with GBDT leaves was 0.69588 when we tested GBDT leaves. And weighted average of them scored 0.6971.So even single model score was not improved, GBDT leaves feature was very important.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 157225,
      "author_name": "Gabriel Moreira",
      "author_url": "",
      "post_date": "2017-01-19T18:49:26.890000",
      "content": "<p>@CuteChibiko I also tried the 3 idiots approach to add GBDT leaves nodes as features to FFM, but did not succeeded improving the score using standard LibFFM, (actually they made my FFM CV score worse). I tried 30 trees with 64 and 128 leaves, using XGBoost and LightGBM.\nDo you think my GBDT model may not have being accurate enought (CV=0.670), while my FFM model was (0.679)?\nOther aspect is that I used a fixed validation set following test set days distribution. So, I trained GBDT on full train set and used the model to predict both the train set and validation set. Maybe that could cause some overfit on FFM model, even with using leaf nodes?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 157180,
      "author_name": "Sameh Faidi",
      "author_url": "",
      "post_date": "2017-01-19T15:18:38.207000",
      "content": "<p>\"impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.\"</p>\n\n<p>what do you mean by impression counts?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 157171,
      "author_name": "Sangxia",
      "author_url": "",
      "post_date": "2017-01-19T14:40:51.577000",
      "content": "<p>I have 2 questions: </p>\n\n<ol>\n<li>Could you explain what XGB leaf feature is?  </li>\n<li>Do you normalize the confidence vectors for documents?</li>\n</ol>\n\n<p>Thanks and congrats!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 159127,
      "author_name": "Gabriel Moreira",
      "author_url": "",
      "post_date": "2017-02-01T00:06:25.283000",
      "content": "<p>Good job @CuteChibiko - I took a look in your implementation and saw it really deserved a try. I've run today a benchmark using some of my datasets in FFM format comparing standard LibFFM and your fork with custom pairwise rank customization. <br>\nI just run the models with the same FFM hyperparameters (without tunning for MAP ranking) and weight 1 for all samples. The environment was a Google Compute Engine instance with 32 CPUs, 256 GB RAM and FFM data in a SSD disk.  </p>\n\n<p>1) Dataset with categorical fields and some binned numeric fields <br>\nParameters: -l 0.0001 -k 6 -t 20 -r 0.2 -s 28 --auto-stop <br>\nRounds: 20  </p>\n\n<p>LibFFM - Cross-validation MAP             = 0.65911 <br>\nLibFFM - Cross-validation MAP + leak = 0.67795 <br>\nLibFFM - Training time                           = 291m  </p>\n\n<p>LibFFM pairwise - Cross-validation MAP             = 0.66040 <br>\nLibFFM pairwise - Cross-validation MAP + leak = 0.67792 <br>\nLibFFM pairwise - Training time                           = 631m  </p>\n\n<hr>\n\n<p>2) Dataset with only categorical fields\nParameters: -l 0.00002 -k 6 -t 50 -r 0.1 -s 28 --auto-stop</p>\n\n<p>LibFFM - Training rounds = 13  (auto-stop on validation logloss) <br>\nLibFFM - Cross-validation MAP             = 0.66428 <br>\nLibFFM - Cross-validation MAP + leak = 0.67859 <br>\nLibFFM - Training time                           = 27m</p>\n\n<p>LibFFM - Training rounds = 26 (auto-stop on validation MAP) <br>\nLibFFM pairwise - Cross-validation MAP             = 0.66552 <br>\nLibFFM pairwise - Cross-validation MAP + leak = 0.67973 <br>\nLibFFM pairwise - Training time                           = 211m</p>\n\n<hr>\n\n<p>I noticed a diff of around 0.001 in CV MAP by for two datasets. Was this the order of difference you could find in your experiments with your FFM with ranking optimization? <br>\nPs. I also noticed a great increase in training logloss (from 0.41009 to 0.56855), but this is expected as the optimization is now driven by MAP and not logloss), right?  </p>\n\n<p>I am worried only about accuracy at this time, but I could notice that training time of your implementation is very slow compared to standard FFM, and that during most of the processing it is using only 1 CPU. Is that because you had not time to parallelize MAP evaluation?   </p>\n\n<p>Which sample weight strategy have you used? More weight on recent samples / clicked samples / display_ids with more ads?  </p>\n\n<p>Thanks again!  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 158836,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2017-01-30T11:06:15.573000",
      "content": "<p>@Gabriel Moreira,</p>\n\n<p>I uploaded this source code. You can find pairwise rank customization <a href=\"https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L443\">around here</a>.</p>\n\n<pre><code>if(param.use_default_train){\n  # default libffm training\n} else {\n  # pairwise rank training\n}\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 157187,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2017-01-19T15:40:34.513000",
      "content": "<p>@Sameh Faidi,</p>\n\n<p>I meant impression count as how many times the ad is shown.\nExplanation of impression in the context of online advertising can be found <a href=\"https://en.wikipedia.org/wiki/Impression_(online_media)\">wikipedia</a>. </p>\n\n<p>For example, impression counts for advertiser_id means how many times ads from the advertiser is shown.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 157179,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2017-01-19T15:10:08.563000",
      "content": "<p>Thank you, Sangzia,</p>\n\n<ol>\n<li>Page 9 of <a href=\"http://www.csie.ntu.edu.tw/~r01922136/kaggle-2014-criteo.pdf\">this presentation by 3 idiots</a> is very good explanation of XGB leaf feature.  </li>\n<li>Document confidence vectors are not normalized, but user confidence vectors are normalized.</li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 159411,
      "author_name": "nomo",
      "author_url": "",
      "post_date": "2017-02-02T05:57:11.737000",
      "content": "<p>Thanks for your kindly share@CuteChibiko . It is really impressive and interesting. I have several questoins about gbdt leaf node as feature.</p>\n\n<p>1). Does the gbdt feature contain label information? like click/ctr. Or if not contain, like impression/pageview count etc.</p>\n\n<p>2). Does the gbdt training data have overlap with ffm training data, If yes, does it suffer from overfitting at first due to label twice used and do you use some method to avoid overfitting?</p>\n\n<p>3). Your weighted average with gbdt leaf's gain is large. Have you compared with ffm containing the gbdt's count feature in some transformation like log without gbdt leafnode, or use gbdt and ffm pctr prediction score together feeded into stacking model. I'm curious that the gain is from gbdt leaf node transformation method or the count/statistics feature itself.</p>\n\n<p>4). Ususallly gbdt training has several hundred or thousand trees, and the cost is large for ffm. How many trees you have used and do you have some hyper-paramter or feature tuning for the less tree numbers?</p>\n\n<p>Thanks a lot.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 159463,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2017-02-02T11:35:06.360000",
          "content": "<p>@nomo,</p>\n\n<p>1) The gbdt feature dose not contain label information but contains impression/pageview count. This is almost same as what I listed in the first post of this topic.</p>\n\n<p>2) Yes, it dose. The gbdt training data was same as ffm training data. But it did not suffer from overfitting.</p>\n\n<p>3) I'm not sure I could understand your question correctly, but we only tried xgb leafnode information itself as feature and did not compared with any variations.</p>\n\n<p>4) We used ntrees=39 and maxDepth=15. Almost all features and hyper-paramters of this model are the same as values of best xgb single model. But only eta was tuned for this model to reduce ntrees.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 159490,
          "author_name": "nomo",
          "author_url": "",
          "post_date": "2017-02-02T14:44:01.707000",
          "content": "<p>@CuteChibiko thank you very much for your detailed explain. It is interesting:-)</p>\n\n<p>I have also tried split data as present/future for xgb stacking training, but the result is neutral to not split one, and am interested in the different and your success trying reason.</p>\n\n<p>1). For a.single model ffm / b.single model xgb / c. stacking model, your present/future split has gains on all of these three models or some of these models?</p>\n\n<p>2). How much gain do your solution has compared with not split modeling? Is it mainly from present part or future part? </p>\n\n<p>3). The future model process is interesting. Does the sample weight use exponent decay, like DecayFactor ^ (LastDay - <strong>SampleDay</strong>) as sample weight?</p>\n\n<p>4). I have also used timestamp feature for split modeling present part, but use same hyper-parameter for the present/future part, and the result is neutral. Do your different parameters for these two part influence a lot? Or some other feature or model method also make a difference, like weighted average of present/future split modeling and not split model.</p>\n\n<p>thanks a lot!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 159552,
          "author_name": "tito",
          "author_url": "",
          "post_date": "2017-02-02T18:48:33.603000",
          "content": "<p>@nomo,</p>\n\n<p>1) Splitting data worked for a and b. And It seemed to work for validation data set of stacking model, but LB score decreased. We did not have enough time to find the reason, so we did not use splitting for stacking model.</p>\n\n<p>2) The single model gain is 0.001 for ffm future model with sample weight, 0.002 for ffm present model with timestamp feature and 0.002 for xgb present model with timestamp.</p>\n\n<p>3) Exactly. We used this formula: </p>\n\n<pre><code>sample weight = 1 + 2 * 0.5 ^ ((max_timestamp - timestamp)/86400000.0)\n</code></pre>\n\n<p>4) Timestamp feature for present data made big gain without any special method. We used almost same hyper-parameters for split model. Only tuned hyper-parameter is iterations/ntrees.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 159137,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2017-02-01T01:45:28.363000",
      "content": "<p>@Gabriel Moreira, thank you for your feedback.</p>\n\n<p>The difference was 0.0015 ~ 0.002 for our models.\nRelatively bigger ETA for LibFFM pairwise made better results.</p>\n\n<p>You are right. It is using only 1 CPU to calculate validation score.\nI don't remember why I commented out <a href=\"https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L500\">here</a> and did not use omp. :(</p>\n\n<p>We added higher weights to recent samples for ‘future’ model. This improved 'future' model score about 0.001.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 158778,
      "author_name": "Gabriel Moreira",
      "author_url": "",
      "post_date": "2017-01-30T01:48:11.027000",
      "content": "<p>Thanks for sharing @CuteChibiko ! You've told you used a customized libffm supporting pairwise rank.\nI am really interested in how the FFM optimization works when using pairwise ranking for libffm.... could you share something about that?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 157575,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-22T08:01:42.010000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "157166": "Congratulations to code monkey, brain-afk and everyone!\r\nAnd I’d like to thank my teammates carl and Little Boat. I could learn lots of things from them.\r\n\r\nAs [Little Boat explained](https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27897/congrats-and-solution-sharing?forumMessageId=157062#post157062), our model is ffm models + xgboost models as first layer, and xgboost as second layer.\r\n\r\nI would like to explain our first layer models here.\r\n\r\n###**Feature extraction**\r\n\r\nFollowing is the list of unobvious features we used.\r\nWe made lots of other features which turned out to be useless.\r\nBut analyzing large scale user behavior was very interesting to me.\r\n\r\n - page view counts of each user.\r\n - page view counts of ad landing page.\r\n - impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.\r\n - landing page confidence vector, where confidence vector is defined as a vector which element is composed of confidence level from documents_*.csv. This feature is used only for ffm as numeric data.\r\n - user confidence vector - this is the average of document confidence vector which is viewed by each user. This feature is used only for ffm as numeric data.\r\n - inner dot product of document confidence vectors of ad impression page and ad landing page .\r\n - inner dot product of user confidence vector and ad landing page document confidence vector\r\n - XGB leaf for ffm feature.\r\n - immediate document viewed after click event.\r\n\r\n###**Modeling and Training**\r\n\r\n####**FFM**\r\nIn addition to [carl’s light-ffm](https://www.kaggle.com/c/outbrain-click-prediction/forums/t/27892/introducing-light-ffm-and-stack-nn), we used another customized libffm which has pair-wise rank and can take sample weight.\r\nBest public LB score is 0.6974.\r\n\r\n####**XGB**\r\n We used xgboost with rank:pairwise objective.\r\nBest public LB score is 0.6885.\r\n\r\n####**Leak Row Excluding**\r\nWe trained ffm models without leak row by removing them from training data first, then blend/merge leak information later.\r\nThis improved single model score.\r\n\r\n####**Present/Future split**\r\nWe split test data into ‘present’ data (in time samples) and ‘future’ data (out time samples), then built models for each data.\r\nHere, Features and hyper parameters for each model are optimized independently.\r\nFor example we used timestamp feature for ‘present’ model, but did not used it for ‘future’ model. And we used sample weight of training data for ‘future’ model to add higher weight to last day than first day.\r\n\r\nPlease feel free to ask any questions.\r\n\r\n",
    "158945": "Great @CuteChibiko ! Thanks for sharing!",
    "157578": "Thank you Yuriy Makarov,\r\n\r\nYes, we tried CTR with some impression threshold. If impression threshold was low, CTR caused overfit. And if impression threshold was high, CTR did not contribute to score. So We did not use CTR in the end.\r\n\r\n\"immediate document viewed after click event.\" means the first document that was seen by a user after the user clicked an ad.  ",
    "157274": "@Gabriel Moreira,\r\n\r\nYour CV score drop, from CV=0.679 to 0.670, is very big, so I think something should be happening.\r\nI'm not sure why, but maybe you are right. We made XGB models for test and validation respectively, so this may cause the difference.\r\n\r\nBTW, our best single score did not improved with GBDT leaves, but it worked for ensemble model.\r\n\r\nOur best single CV score without GBDT leaves was 0.69586 and score with GBDT leaves was 0.69588 when we tested GBDT leaves. And weighted average of them scored 0.6971.So even single model score was not improved, GBDT leaves feature was very important.",
    "157225": "@CuteChibiko I also tried the 3 idiots approach to add GBDT leaves nodes as features to FFM, but did not succeeded improving the score using standard LibFFM, (actually they made my FFM CV score worse). I tried 30 trees with 64 and 128 leaves, using XGBoost and LightGBM.\r\nDo you think my GBDT model may not have being accurate enought (CV=0.670), while my FFM model was (0.679)?\r\nOther aspect is that I used a fixed validation set following test set days distribution. So, I trained GBDT on full train set and used the model to predict both the train set and validation set. Maybe that could cause some overfit on FFM model, even with using leaf nodes?",
    "157180": "\"impression counts for each ad_id, landing document_id, campaign_id and advertiser_id.\"\r\n\r\nwhat do you mean by impression counts?",
    "157171": "I have 2 questions: \r\n\r\n 1. Could you explain what XGB leaf feature is?  \r\n 2. Do you normalize the confidence vectors for documents?\r\n\r\nThanks and congrats!",
    "159127": "Good job @CuteChibiko - I took a look in your implementation and saw it really deserved a try. I've run today a benchmark using some of my datasets in FFM format comparing standard LibFFM and your fork with custom pairwise rank customization.    \r\nI just run the models with the same FFM hyperparameters (without tunning for MAP ranking) and weight 1 for all samples. The environment was a Google Compute Engine instance with 32 CPUs, 256 GB RAM and FFM data in a SSD disk.  \r\n\r\n1) Dataset with categorical fields and some binned numeric fields  \r\nParameters: -l 0.0001 -k 6 -t 20 -r 0.2 -s 28 --auto-stop  \r\nRounds: 20  \r\n  \r\nLibFFM - Cross-validation MAP             = 0.65911  \r\nLibFFM - Cross-validation MAP + leak = 0.67795  \r\nLibFFM - Training time                           = 291m  \r\n  \r\nLibFFM pairwise - Cross-validation MAP             = 0.66040  \r\nLibFFM pairwise - Cross-validation MAP + leak = 0.67792  \r\nLibFFM pairwise - Training time                           = 631m  \r\n\r\n----------------------------------------------------\r\n\r\n2) Dataset with only categorical fields\r\nParameters: -l 0.00002 -k 6 -t 50 -r 0.1 -s 28 --auto-stop\r\n\r\nLibFFM - Training rounds = 13  (auto-stop on validation logloss)  \r\nLibFFM - Cross-validation MAP             = 0.66428  \r\nLibFFM - Cross-validation MAP + leak = 0.67859  \r\nLibFFM - Training time                           = 27m\r\n\r\nLibFFM - Training rounds = 26 (auto-stop on validation MAP)  \r\nLibFFM pairwise - Cross-validation MAP             = 0.66552  \r\nLibFFM pairwise - Cross-validation MAP + leak = 0.67973  \r\nLibFFM pairwise - Training time                           = 211m\r\n\r\n-----------------------------------\r\n  \r\nI noticed a diff of around 0.001 in CV MAP by for two datasets. Was this the order of difference you could find in your experiments with your FFM with ranking optimization?   \r\nPs. I also noticed a great increase in training logloss (from 0.41009 to 0.56855), but this is expected as the optimization is now driven by MAP and not logloss), right?  \r\n   \r\nI am worried only about accuracy at this time, but I could notice that training time of your implementation is very slow compared to standard FFM, and that during most of the processing it is using only 1 CPU. Is that because you had not time to parallelize MAP evaluation?   \r\n\r\nWhich sample weight strategy have you used? More weight on recent samples / clicked samples / display_ids with more ads?  \r\n\r\nThanks again!  ",
    "158836": "@Gabriel Moreira,\r\n\r\nI uploaded this source code. You can find pairwise rank customization [around here](https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L443).\r\n\r\n    if(param.use_default_train){\r\n      # default libffm training\r\n    } else {\r\n      # pairwise rank training\r\n    }",
    "157187": "@Sameh Faidi,\r\n\r\nI meant impression count as how many times the ad is shown.\r\nExplanation of impression in the context of online advertising can be found [wikipedia](https://en.wikipedia.org/wiki/Impression_(online_media)). \r\n\r\nFor example, impression counts for advertiser_id means how many times ads from the advertiser is shown.\r\n",
    "157179": "Thank you, Sangzia,\r\n\r\n 1. Page 9 of [this presentation by 3 idiots](http://www.csie.ntu.edu.tw/~r01922136/kaggle-2014-criteo.pdf) is very good explanation of XGB leaf feature.  \r\n 2. Document confidence vectors are not normalized, but user confidence vectors are normalized.",
    "159411": "Thanks for your kindly share@CuteChibiko . It is really impressive and interesting. I have several questoins about gbdt leaf node as feature.\n\n1). Does the gbdt feature contain label information? like click/ctr. Or if not contain, like impression/pageview count etc.\n\n2). Does the gbdt training data have overlap with ffm training data, If yes, does it suffer from overfitting at first due to label twice used and do you use some method to avoid overfitting?\n\n3). Your weighted average with gbdt leaf's gain is large. Have you compared with ffm containing the gbdt's count feature in some transformation like log without gbdt leafnode, or use gbdt and ffm pctr prediction score together feeded into stacking model. I'm curious that the gain is from gbdt leaf node transformation method or the count/statistics feature itself.\n\n4). Ususallly gbdt training has several hundred or thousand trees, and the cost is large for ffm. How many trees you have used and do you have some hyper-paramter or feature tuning for the less tree numbers?\n\n\nThanks a lot.\n",
    "159137": "@Gabriel Moreira, thank you for your feedback.\r\n\r\nThe difference was 0.0015 ~ 0.002 for our models.\r\nRelatively bigger ETA for LibFFM pairwise made better results.\r\n\r\nYou are right. It is using only 1 CPU to calculate validation score.\r\nI don't remember why I commented out [here](https://github.com/CuteChibiko/Outbrain-Click-Prediction/blob/master/libffm_pairwise/ffm.cpp#L500) and did not use omp. :(\r\n\r\nWe added higher weights to recent samples for ‘future’ model. This improved 'future' model score about 0.001.",
    "158778": "Thanks for sharing @CuteChibiko ! You've told you used a customized libffm supporting pairwise rank.\r\nI am really interested in how the FFM optimization works when using pairwise ranking for libffm.... could you share something about that?",
    "157575": ""
  }
}