{
  "id": 27897,
  "title": "Congrats and Solution Sharing",
  "url": "/competitions/outbrain-click-prediction/writeups/freshdesk-ml-congrats-and-solution-sharing",
  "author_name": "",
  "post_date": "2017-01-19T02:08:09.847Z",
  "votes": 15,
  "comment_count": 10,
  "views": 898,
  "content": "<p>Congratulations code monkey, brain-afk and Three Data Points for the top 3 finish. Congratulations to other top finishers as well. </p>\n\n<p>It was a very interesting competition due to lot of factors such as Data size, number of tables to use, test set having both in time and out time samples, ensembling for map metric etc</p>\n\n<p>It will be really helpful for other people in the community if the top finishers share their ideas, codes and the thinking that went behind them in this thread.</p>\n\n<p>Thank you Kaggle and outbrain for this nice competition. Thanks to my team mates and everyone else for making this competition more lively.</p>\n\n<p>Thank you.! </p>",
  "messages": [
    {
      "id": "157058",
      "postDate": "01/19/2017 02:08:09",
      "content": "<p>Congratulations code monkey, brain-afk and Three Data Points for the top 3 finish. Congratulations to other top finishers as well. </p>\n\n<p>It was a very interesting competition due to lot of factors such as Data size, number of tables to use, test set having both in time and out time samples, ensembling for map metric etc</p>\n\n<p>It will be really helpful for other people in the community if the top finishers share their ideas, codes and the thinking that went behind them in this thread.</p>\n\n<p>Thank you Kaggle and outbrain for this nice competition. Thanks to my team mates and everyone else for making this competition more lively.</p>\n\n<p>Thank you.! </p>",
      "rawMarkdown": "Congratulations code monkey, brain-afk and Three Data Points for the top 3 finish. Congratulations to other top finishers as well. \r\n\r\nIt was a very interesting competition due to lot of factors such as Data size, number of tables to use, test set having both in time and out time samples, ensembling for map metric etc\r\n\r\nIt will be really helpful for other people in the community if the top finishers share their ideas, codes and the thinking that went behind them in this thread.\r\n\r\nThank you Kaggle and outbrain for this nice competition. Thanks to my team mates and everyone else for making this competition more lively.\r\n\r\nThank you.!",
      "votes": null
    },
    {
      "id": "157062",
      "postDate": "01/19/2017 02:28:50",
      "content": "<p>Big congrats to the top 2 finishes!! Both did amazing! And also huge congrats to Andrii Cherednychenko for getting the score playing solo.  I was actually thinking about merging with you on the last day of merging deadline.</p>\n\n<p>Carl and CuteChibiko should actually be given all the credits to. I just did some trivial stuff and tried and failed a bunch of ideas.</p>\n\n<p>So I will briefly describe our approaches and let Carl and CuteChibiko dive into details later.</p>\n\n<p>our final approach in one sentence:\nffm models + xgboost models as first layer, and xgboost as second layer</p>\n\n<p>for ffm models, we have two loss functions, softmax loss (that distributed within each group) and pairwise rank loss.\nfor xgboost models, it is just pairwise rank.</p>\n\n<p>I was focusing on ftrl models for a while at the beginning of the competition. So I tried lambda rank and pairwise rank with them. If I select two way interactions carefully, the ftrl model performance is actually not too far away from ffm model but it can take more than 10+ hours to generate and it contributed little to second layer ensemble, so I moved on.</p>\n\n<p>In the last week, I was focusing on nn models. It gets worse result than ffm model but about the same as xgboost model. However one single run of nn models doesn't really give a very stable result (even with 50m data points!!), and I didn't really have time for generating say, 10 - 20 nn models and then average them, so I gave up there too.</p>",
      "rawMarkdown": "Big congrats to the top 2 finishes!! Both did amazing! And also huge congrats to Andrii Cherednychenko for getting the score playing solo.  I was actually thinking about merging with you on the last day of merging deadline.\r\n\r\nCarl and CuteChibiko should actually be given all the credits to. I just did some trivial stuff and tried and failed a bunch of ideas.\r\n\r\nSo I will briefly describe our approaches and let Carl and CuteChibiko dive into details later.\r\n\r\nour final approach in one sentence:\r\nffm models + xgboost models as first layer, and xgboost as second layer\r\n\r\nfor ffm models, we have two loss functions, softmax loss (that distributed within each group) and pairwise rank loss.\r\nfor xgboost models, it is just pairwise rank.\r\n\r\nI was focusing on ftrl models for a while at the beginning of the competition. So I tried lambda rank and pairwise rank with them. If I select two way interactions carefully, the ftrl model performance is actually not too far away from ffm model but it can take more than 10+ hours to generate and it contributed little to second layer ensemble, so I moved on.\r\n\r\nIn the last week, I was focusing on nn models. It gets worse result than ffm model but about the same as xgboost model. However one single run of nn models doesn't really give a very stable result (even with 50m data points!!), and I didn't really have time for generating say, 10 - 20 nn models and then average them, so I gave up there too.",
      "votes": null
    },
    {
      "id": "157107",
      "postDate": "01/19/2017 08:30:14",
      "content": "<p>i also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. </p>\n\n<p>for the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. </p>\n\n<p>i also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.</p>",
      "rawMarkdown": "i also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. \r\n\r\nfor the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. \r\n\r\ni also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.",
      "votes": null
    },
    {
      "id": "157115",
      "postDate": "01/19/2017 09:21:06",
      "content": "<p>Congrats to the winners, and well done to all participants. :-)</p>\n\n<p>My solution is a simple average of 3 diversified xgboost models. </p>\n\n<p><strong>Primary features</strong></p>\n\n<ul>\n<li>Calculation of click ratio for categorical features</li>\n<li>Count of ads per display_id</li>\n<li>Count of page_views for ad-related categoricals</li>\n<li>Bucketing of numerical features</li>\n</ul>\n\n<p>Best xgboost\nPB: 0.68461</p>\n\n<p>The models contained between 40-60 features and were trained on 90% data, with the remaining 10% as validation.</p>\n\n<p><strong>Best models parameters</strong></p>\n\n<ul>\n<li>Metric: logloss</li>\n<li>Colsample: 0.30</li>\n<li>Subsample: 0.85</li>\n<li>Depth: 7</li>\n<li>ETA: 0.1 (or less, can't remember)</li>\n<li>Child weight: 5</li>\n<li>Gamma: 0</li>\n<li>Rounds: ~5.000</li>\n</ul>",
      "rawMarkdown": "Congrats to the winners, and well done to all participants. :-)\r\n\r\nMy solution is a simple average of 3 diversified xgboost models. \r\n\r\n**Primary features**\r\n\r\n- Calculation of click ratio for categorical features\r\n- Count of ads per display_id\r\n- Count of page_views for ad-related categoricals\r\n- Bucketing of numerical features\r\n\r\n\r\nBest xgboost\r\nPB: 0.68461\r\n\r\nThe models contained between 40-60 features and were trained on 90% data, with the remaining 10% as validation.\r\n\r\n**Best models parameters**\r\n\r\n- Metric: logloss\r\n- Colsample: 0.30\r\n- Subsample: 0.85\r\n- Depth: 7\r\n- ETA: 0.1 (or less, can't remember)\r\n- Child weight: 5\r\n- Gamma: 0\r\n- Rounds: ~5.000",
      "votes": null
    },
    {
      "id": "157119",
      "postDate": "01/19/2017 09:43:33",
      "content": "<p>[quote=Sameh Faidi;157107]</p>\n\n<p>i also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. </p>\n\n<p>for the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. </p>\n\n<p>i also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.</p>\n\n<p>[/quote]</p>\n\n<p>Could you explain feature (1)? By ad landing page do you mean the page where the context is displayed? If I remember correctly, for almost all contexts one could find such user-page record in page views, is that what you saw as well?</p>\n\n<p>The time features I created didn't help with my fm models, but was useful in my second level xgboost model.</p>",
      "rawMarkdown": "[quote=Sameh Faidi;157107]\r\n\r\ni also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. \r\n\r\nfor the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. \r\n\r\ni also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.\r\n\r\n[/quote]\r\n\r\nCould you explain feature (1)? By ad landing page do you mean the page where the context is displayed? If I remember correctly, for almost all contexts one could find such user-page record in page views, is that what you saw as well?\r\n\r\nThe time features I created didn't help with my fm models, but was useful in my second level xgboost model.",
      "votes": null
    },
    {
      "id": "157168",
      "postDate": "01/19/2017 14:28:23",
      "content": "<p>@Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?</p>",
      "rawMarkdown": "Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?",
      "votes": null
    },
    {
      "id": "157182",
      "postDate": "01/19/2017 15:19:27",
      "content": "<p>[quote=FengLi;157168]</p>\n\n<p>@Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?</p>\n\n<p>[/quote]</p>\n\n<p>The approximately 5.000 rounds were determined by auto stopping. (10% data used for validation)</p>\n\n<p>Among the categorical variables with multiple values, I selected the single variable with the highest confidence level, and added the confidence level as a bucketed variable. Any NaN's are set to \"-1\" (out of range). </p>",
      "rawMarkdown": "[quote=FengLi;157168]\r\n\r\n@Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?\r\n\r\n[/quote]\r\n\r\nThe approximately 5.000 rounds were determined by auto stopping. (10% data used for validation)\r\n\r\nAmong the categorical variables with multiple values, I selected the single variable with the highest confidence level, and added the confidence level as a bucketed variable. Any NaN's are set to \"-1\" (out of range).",
      "votes": null
    },
    {
      "id": "157204",
      "postDate": "01/19/2017 17:06:36",
      "content": "<p>5000 rounds is impressive! How long does it take to train your model and what resources did you use?</p>",
      "rawMarkdown": "5000 rounds is impressive! How long does it take to train your model and what resources did you use?",
      "votes": null
    },
    {
      "id": "157215",
      "postDate": "01/19/2017 17:59:44",
      "content": "<p>This Kaggle competition was special for me, as the first I was really engaged on. Three months with lots of learning and little sleep. To be in the first LB page scroll (19th) was rewarding for me.\nHere is my detailed solution (questions and feedbacks are welcome):</p>\n\n<p>Ps. More details about my approach in this <a href=\"https://medium.com/unstructured/how-feature-engineering-can-help-you-do-well-in-a-kaggle-competition-part-i-9cc9a883514d\">post series</a>.</p>\n\n<h3>Feature Engineering</h3>\n\n<p>All feature engineering made using a little PySpark cluster (Python). An example of the features of Spark SQL may be found on this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\">EDA</a>.</p>\n\n<p><em><strong>User profile</strong></em>\n(based on page_views)</p>\n\n<ul>\n<li>User has previously viewed the ad document? (page_views)</li>\n<li>User page views count</li>\n<li>Categories, Topics and Entities of the documents users have previously viewed (weighted by confidence and TF-IDF) , to model users preferences in a Content-Based Filtering approach. <br>\nPs. According to this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\">EDA Kernel</a>, about 75% of users in events dataset had at least one additional view on page views dataset (besides the clicks logged as events)</li>\n</ul>\n\n<p><em><strong>Documents and Ads</strong></em></p>\n\n<ul>\n<li>Days elapsed since Ad document and Event (landing page) document were published</li>\n<li>Avg page views by distinct users</li>\n<li>Documents and Ads views count</li>\n</ul>\n\n<p><em><strong>Events</strong></em></p>\n\n<ul>\n<li>Event_hour, adjusted for local timezone (based on state geolocation) and binned (morning, afternoon, ...)</li>\n<li>Is_weekend?</li>\n</ul>\n\n<p><em><strong>Categorical fields</strong></em>\nOne-Hot encoding of categorical fields (about resulting 126K features)</p>\n\n<ul>\n<li>ad_id   </li>\n<li>doc_event_id</li>\n<li>doc_ad_id</li>\n<li>ad_advertiser</li>\n<li>doc_ad_category_id</li>\n<li>doc_ad_entity_id</li>\n<li>doc_ad_publisher_id</li>\n<li>doc_ad_source_id</li>\n<li>doc_ad_topic_id</li>\n<li>doc_event_category_id</li>\n<li>doc_event_entity_id</li>\n<li>doc_event_publisher_id</li>\n<li>doc_event_source_id</li>\n<li>doc_event_topic_id</li>\n<li>event_country</li>\n<li>event_country_state</li>\n<li>event_geo_location</li>\n<li>event_platform</li>\n<li>traffic_source</li>\n</ul>\n\n<p><em><strong>Avg CTR</strong></em> <br>\nAverage CTR (#clicks / #views)  based on categorical fields combinations </p>\n\n<ul>\n<li>ad_id  </li>\n<li>document_id  </li>\n<li>publisher_id  </li>\n<li>advertiser_id  </li>\n<li>campain_id  </li>\n<li>doc_event + doc_ad  </li>\n<li>source_id  </li>\n<li>source_id + country  </li>\n<li>entity_id  </li>\n<li>entity_id + country  </li>\n<li>topic_id  </li>\n<li>topic_id + country  </li>\n<li>category_id  </li>\n<li>category_id + country  </li>\n</ul>\n\n<p><em><strong>Content-Based Similarities</strong></em> <br>\nCosine similarity between user profile and ad doc aspects vectors (TF-IDF)  </p>\n\n<ul>\n<li>user_doc_ad_sim_categories  </li>\n<li>user_doc_ad_sim_topics  </li>\n<li>user_doc_ad_sim_entities  </li>\n</ul>\n\n<p>Cosine similarity between event doc (landing page) and ad doc aspects vectors (TF-IDF)</p>\n\n<ul>\n<li>doc_event_doc_ad_sim_categories</li>\n<li>doc_event_doc_ad_sim_topics</li>\n<li>doc_event_doc_ad_sim_entities</li>\n</ul>\n\n<h3>Cross-validation</h3>\n\n<p>I used a fixed validation set with the same days distribution of the test set (20% of clicks of training events in the first 11 days and all events in days 12 and 13). The alignment of my CV and public LB was accurate at the 4th decimal digit. </p>\n\n<h3>1st level models</h3>\n\n<ul>\n<li>LibFFM (hashed categorical fields) - 0.67932</li>\n<li>LibFFM (hashed categorical fields and binned numeric features) - 0.6784</li>\n<li>LibFFM (hashed categorical fields and binned numeric features trained with the latter 30% events (last days) - 0.6736</li>\n<li>FTRL (hashed categorical quadratic interactions) - 0.67659</li>\n<li>VW (-ftrl, hashed quad interactions among categorical + numeric binned) - 0.6751</li>\n<li>VW (-ftrl, hashed quad interactions only among categorical + numeric raw (not binned)) - 0.67691</li>\n<li>LightGBM (only numeric features) - 0.6707</li>\n<li>XGBoost (numeric and OHE categories) - 0.6689</li>\n<li>RankLib (Random Forests, LambdaMART, ListNet, AdaBoost, RankBoost, Coordinate Ascent) (only numeric fieds)  ~ 0.65-0.66 each model</li>\n</ul>\n\n<h3>2nd level (emsembling)</h3>\n\n<p>I am newbie in emsembling and had little time to explore it. <br>\nMy first approach was to try many kinds of weighted averages (arithmetic, geometric, harmonic) of model predictions. The best setting was a weighted average of inverse logarithmic from the top 4 models (LB=0.68418).</p>\n\n<p>My next approach was to apply XGBoost with rank objective, which is a natural choice as it optimizes the contest metric (MAP), to emsemble models predictions and some numeric features. \nAs I worked with a fixed validation set in 1st level, I had to train model only using validation set (blending) model predictions and original features.  So, I split validation set in train and eval set (50% each), where the eval also followed the test set days distribution. I excluded leaked rows for training, to let models learn other aspects as leak was free :). Using my 6 top models predictions and 15 numeric features I got my best LB score yesterday: <em>0.68716</em>.</p>\n\n<p>I also trained 9 FFM models on subsets of the most frequent event countries (US, CA, GB, AU, Other) and states (CA, FL, TX, NY), and added those predictions and equivalente OHE geo categories in the emsemble, but could not finish before the competition end. I will try an emsemble with all my models in the next few days to see how it would have been.</p>\n\n<p>Kaggle is really a university for cutting edge Machine Learning.</p>\n\n<h3>What didn't work</h3>\n\n<ul>\n<li>Spark ALS Matrix Factorization - User-Based Collaborative Filtering with rank=10 (the maximum I could use in a Spark cluster) underperformed (MAP=0.56). Probably because of sparseness of data.</li>\n<li>One-Hot Encoding Categorical Features - I used OHE as input for XGBoost, but according to better performing models of other competitors using raw id of categorical fields would be fine for trees emsembles. OHE are costly to build, because they require lots of dictionaries. And FFM, FTRL and VW worked well with hashed features.</li>\n<li>RankLib is a collection of algorithms for ranking, using the same input format, which in nice. Therefore, their Java implementation use lots of memory, so I could run on just a sample o numeric data. And their predictions underperformed compared to XGBoost and LightGBM.</li>\n<li>Using GBDT leaf nodes as features to FFM (winning approach in Criteo competition, by 3 idiots) did not worked for me. They actually reduced by CV score.</li>\n<li>Keeping a fixed validation set using the same test set days distribution allowed a CV very aligned with LB. Therefore, in the 2nd level (emsemble), I had to train model using only validation set (blending) model predictions and engineered features. This blending was my best approach, but studying in the forums, I saw that it could be better to run out-of-fold predictions for fixed folds in the 1st level. This way, I would be able to use the full training set, stacking in the 2nd level model predictions (without leak) together with engineered features.</li>\n</ul>",
      "rawMarkdown": "This Kaggle competition was special for me, as the first I was really engaged on. Three months with lots of learning and little sleep. To be in the first LB page scroll (19th) was rewarding for me.\nHere is my detailed solution (questions and feedbacks are welcome):\n\nPs. More details about my approach in this [post series][1].\n\n### Feature Engineering \n\nAll feature engineering made using a little PySpark cluster (Python). An example of the features of Spark SQL may be found on this [EDA][2].\n\n***User profile***\n(based on page_views)\n\n - User has previously viewed the ad document? (page_views)\n - User page views count\n - Categories, Topics and Entities of the documents users have previously viewed (weighted by confidence and TF-IDF) , to model users preferences in a Content-Based Filtering approach.  \nPs. According to this [EDA Kernel][3], about 75% of users in events dataset had at least one additional view on page views dataset (besides the clicks logged as events)\n\n***Documents and Ads***\n\n - Days elapsed since Ad document and Event (landing page) document were published\n - Avg page views by distinct users\n - Documents and Ads views count\n\n***Events***\n\n - Event_hour, adjusted for local timezone (based on state geolocation) and binned (morning, afternoon, ...)\n - Is_weekend?\n\n***Categorical fields***\nOne-Hot encoding of categorical fields (about resulting 126K features)\n\n - ad_id   \n - doc_event_id\n - doc_ad_id\n - ad_advertiser\n - doc_ad_category_id\n - doc_ad_entity_id\n - doc_ad_publisher_id\n - doc_ad_source_id\n - doc_ad_topic_id\n - doc_event_category_id\n - doc_event_entity_id\n - doc_event_publisher_id\n - doc_event_source_id\n - doc_event_topic_id\n - event_country\n - event_country_state\n - event_geo_location\n - event_platform\n - traffic_source\n\n***Avg CTR***  \nAverage CTR (\\#clicks / \\#views)  based on categorical fields combinations \n \n - ad_id  \n - document_id  \n - publisher_id  \n - advertiser_id  \n - campain_id  \n - doc_event + doc_ad  \n - source_id  \n - source_id + country  \n - entity_id  \n - entity_id + country  \n - topic_id  \n - topic_id + country  \n - category_id  \n - category_id + country  \n\n***Content-Based Similarities***  \nCosine similarity between user profile and ad doc aspects vectors (TF-IDF)  \n\n - user_doc_ad_sim_categories  \n - user_doc_ad_sim_topics  \n - user_doc_ad_sim_entities  \n\nCosine similarity between event doc (landing page) and ad doc aspects vectors (TF-IDF)\n\n - doc_event_doc_ad_sim_categories\n - doc_event_doc_ad_sim_topics\n - doc_event_doc_ad_sim_entities\n\n### Cross-validation  \n\nI used a fixed validation set with the same days distribution of the test set (20% of clicks of training events in the first 11 days and all events in days 12 and 13). The alignment of my CV and public LB was accurate at the 4th decimal digit. \n\n### 1st level models \n\n - LibFFM (hashed categorical fields) - 0.67932\n - LibFFM (hashed categorical fields and binned numeric features) - 0.6784\n - LibFFM (hashed categorical fields and binned numeric features trained with the latter 30% events (last days) - 0.6736\n - FTRL (hashed categorical quadratic interactions) - 0.67659\n - VW (-ftrl, hashed quad interactions among categorical + numeric binned) - 0.6751\n - VW (-ftrl, hashed quad interactions only among categorical + numeric raw (not binned)) - 0.67691\n - LightGBM (only numeric features) - 0.6707\n - XGBoost (numeric and OHE categories) - 0.6689\n - RankLib (Random Forests, LambdaMART, ListNet, AdaBoost, RankBoost, Coordinate Ascent) (only numeric fieds)  ~ 0.65-0.66 each model\n\n### 2nd level (emsembling)\nI am newbie in emsembling and had little time to explore it.  \nMy first approach was to try many kinds of weighted averages (arithmetic, geometric, harmonic) of model predictions. The best setting was a weighted average of inverse logarithmic from the top 4 models (LB=0.68418).\n\nMy next approach was to apply XGBoost with rank objective, which is a natural choice as it optimizes the contest metric (MAP), to emsemble models predictions and some numeric features. \nAs I worked with a fixed validation set in 1st level, I had to train model only using validation set (blending) model predictions and original features.  So, I split validation set in train and eval set (50% each), where the eval also followed the test set days distribution. I excluded leaked rows for training, to let models learn other aspects as leak was free :). Using my 6 top models predictions and 15 numeric features I got my best LB score yesterday: *0.68716*.\n\nI also trained 9 FFM models on subsets of the most frequent event countries (US, CA, GB, AU, Other) and states (CA, FL, TX, NY), and added those predictions and equivalente OHE geo categories in the emsemble, but could not finish before the competition end. I will try an emsemble with all my models in the next few days to see how it would have been.\n\nKaggle is really a university for cutting edge Machine Learning.\n\n### What didn't work\n- Spark ALS Matrix Factorization - User-Based Collaborative Filtering with rank=10 (the maximum I could use in a Spark cluster) underperformed (MAP=0.56). Probably because of sparseness of data.\n- One-Hot Encoding Categorical Features - I used OHE as input for XGBoost, but according to better performing models of other competitors using raw id of categorical fields would be fine for trees emsembles. OHE are costly to build, because they require lots of dictionaries. And FFM, FTRL and VW worked well with hashed features.\n- RankLib is a collection of algorithms for ranking, using the same input format, which in nice. Therefore, their Java implementation use lots of memory, so I could run on just a sample o numeric data. And their predictions underperformed compared to XGBoost and LightGBM.\n- Using GBDT leaf nodes as features to FFM (winning approach in Criteo competition, by 3 idiots) did not worked for me. They actually reduced by CV score.\n- Keeping a fixed validation set using the same test set days distribution allowed a CV very aligned with LB. Therefore, in the 2nd level (emsemble), I had to train model using only validation set (blending) model predictions and engineered features. This blending was my best approach, but studying in the forums, I saw that it could be better to run out-of-fold predictions for fixed folds in the 1st level. This way, I would be able to use the full training set, stacking in the 2nd level model predictions (without leak) together with engineered features.\n\n\n  [1]: https://medium.com/unstructured/how-feature-engineering-can-help-you-do-well-in-a-kaggle-competition-part-i-9cc9a883514d\n  [2]: https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\n  [3]: https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion",
      "votes": null
    },
    {
      "id": "157322",
      "postDate": "01/20/2017 08:06:48",
      "content": "<p>[quote=sswt;157204]</p>\n\n<p>5000 rounds is impressive! How long does it take to train your model and what resources did you use?</p>\n\n<p>[/quote]</p>\n\n<p>It took 4-5 days on a 6 core CPU.</p>",
      "rawMarkdown": "[quote=sswt;157204]\r\n\r\n5000 rounds is impressive! How long does it take to train your model and what resources did you use?\r\n\r\n[/quote]\r\n\r\nIt took 4-5 days on a 6 core CPU.",
      "votes": null
    },
    {
      "id": "157397",
      "postDate": "01/20/2017 19:20:12",
      "content": "<p>I made available my part of our solution at <a href=\"https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle\">https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle</a>. The README file describes the approach in more details.</p>\n\n<p>In short, my contribution was 5 models on the first level: SVM, FTRL, XGB and ET on mean target value features, FFM on XGB leaves. There were some other models like VW or FFM without leaves, but at the end they didn't contribute much to the ensemble, so they aren't included in the code. \nAlso we had 3 more models from Diaman (I'll let him describe his approach).</p>\n\n<p>The second level model was an XGB with pairwise loss, which we trained only on the half of the data. </p>\n\n<p>Before merging we were around ~40 position, but when we combined the models, we jumped to 15th. After that going up from 15th was really tough. </p>\n\n<p>I learned a lot from this competition, thanks everybody. Also, big thanks to my teammate - it was a lot of fun. </p>\n\n<p>See you all in the next competitions ;-) </p>",
      "rawMarkdown": "I made available my part of our solution at [https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle][1]. The README file describes the approach in more details.\r\n\r\nIn short, my contribution was 5 models on the first level: SVM, FTRL, XGB and ET on mean target value features, FFM on XGB leaves. There were some other models like VW or FFM without leaves, but at the end they didn't contribute much to the ensemble, so they aren't included in the code. \r\nAlso we had 3 more models from Diaman (I'll let him describe his approach).\r\n\r\nThe second level model was an XGB with pairwise loss, which we trained only on the half of the data. \r\n\r\nBefore merging we were around ~40 position, but when we combined the models, we jumped to 15th. After that going up from 15th was really tough. \r\n\r\nI learned a lot from this competition, thanks everybody. Also, big thanks to my teammate - it was a lot of fun. \r\n\r\nSee you all in the next competitions ;-) \r\n\r\n  [1]: https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 157062,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "01/19/2017 02:28:50",
      "content": "<p>Big congrats to the top 2 finishes!! Both did amazing! And also huge congrats to Andrii Cherednychenko for getting the score playing solo.  I was actually thinking about merging with you on the last day of merging deadline.</p>\n\n<p>Carl and CuteChibiko should actually be given all the credits to. I just did some trivial stuff and tried and failed a bunch of ideas.</p>\n\n<p>So I will briefly describe our approaches and let Carl and CuteChibiko dive into details later.</p>\n\n<p>our final approach in one sentence:\nffm models + xgboost models as first layer, and xgboost as second layer</p>\n\n<p>for ffm models, we have two loss functions, softmax loss (that distributed within each group) and pairwise rank loss.\nfor xgboost models, it is just pairwise rank.</p>\n\n<p>I was focusing on ftrl models for a while at the beginning of the competition. So I tried lambda rank and pairwise rank with them. If I select two way interactions carefully, the ftrl model performance is actually not too far away from ffm model but it can take more than 10+ hours to generate and it contributed little to second layer ensemble, so I moved on.</p>\n\n<p>In the last week, I was focusing on nn models. It gets worse result than ffm model but about the same as xgboost model. However one single run of nn models doesn't really give a very stable result (even with 50m data points!!), and I didn't really have time for generating say, 10 - 20 nn models and then average them, so I gave up there too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157107,
      "author_name": "samehif",
      "author_url": "",
      "post_date": "01/19/2017 08:30:14",
      "content": "<p>i also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. </p>\n\n<p>for the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. </p>\n\n<p>i also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 157119,
          "author_name": "sangxia",
          "author_url": "",
          "post_date": "01/19/2017 09:43:33",
          "content": "<p>[quote=Sameh Faidi;157107]</p>\n\n<p>i also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. </p>\n\n<p>for the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. </p>\n\n<p>i also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.</p>\n\n<p>[/quote]</p>\n\n<p>Could you explain feature (1)? By ad landing page do you mean the page where the context is displayed? If I remember correctly, for almost all contexts one could find such user-page record in page views, is that what you saw as well?</p>\n\n<p>The time features I created didn't help with my fm models, but was useful in my second level xgboost model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157115,
      "author_name": "frederik",
      "author_url": "",
      "post_date": "01/19/2017 09:21:06",
      "content": "<p>Congrats to the winners, and well done to all participants. :-)</p>\n\n<p>My solution is a simple average of 3 diversified xgboost models. </p>\n\n<p><strong>Primary features</strong></p>\n\n<ul>\n<li>Calculation of click ratio for categorical features</li>\n<li>Count of ads per display_id</li>\n<li>Count of page_views for ad-related categoricals</li>\n<li>Bucketing of numerical features</li>\n</ul>\n\n<p>Best xgboost\nPB: 0.68461</p>\n\n<p>The models contained between 40-60 features and were trained on 90% data, with the remaining 10% as validation.</p>\n\n<p><strong>Best models parameters</strong></p>\n\n<ul>\n<li>Metric: logloss</li>\n<li>Colsample: 0.30</li>\n<li>Subsample: 0.85</li>\n<li>Depth: 7</li>\n<li>ETA: 0.1 (or less, can't remember)</li>\n<li>Child weight: 5</li>\n<li>Gamma: 0</li>\n<li>Rounds: ~5.000</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157168,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "01/19/2017 14:28:23",
      "content": "<p>@Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?</p>",
      "votes": null,
      "replies": [
        {
          "id": 157182,
          "author_name": "frederik",
          "author_url": "",
          "post_date": "01/19/2017 15:19:27",
          "content": "<p>[quote=FengLi;157168]</p>\n\n<p>@Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?</p>\n\n<p>[/quote]</p>\n\n<p>The approximately 5.000 rounds were determined by auto stopping. (10% data used for validation)</p>\n\n<p>Among the categorical variables with multiple values, I selected the single variable with the highest confidence level, and added the confidence level as a bucketed variable. Any NaN's are set to \"-1\" (out of range). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157204,
      "author_name": "savvok",
      "author_url": "",
      "post_date": "01/19/2017 17:06:36",
      "content": "<p>5000 rounds is impressive! How long does it take to train your model and what resources did you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 157322,
          "author_name": "frederik",
          "author_url": "",
          "post_date": "01/20/2017 08:06:48",
          "content": "<p>[quote=sswt;157204]</p>\n\n<p>5000 rounds is impressive! How long does it take to train your model and what resources did you use?</p>\n\n<p>[/quote]</p>\n\n<p>It took 4-5 days on a 6 core CPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157215,
      "author_name": "gspmoreira",
      "author_url": "",
      "post_date": "01/19/2017 17:59:44",
      "content": "<p>This Kaggle competition was special for me, as the first I was really engaged on. Three months with lots of learning and little sleep. To be in the first LB page scroll (19th) was rewarding for me.\nHere is my detailed solution (questions and feedbacks are welcome):</p>\n\n<p>Ps. More details about my approach in this <a href=\"https://medium.com/unstructured/how-feature-engineering-can-help-you-do-well-in-a-kaggle-competition-part-i-9cc9a883514d\">post series</a>.</p>\n\n<h3>Feature Engineering</h3>\n\n<p>All feature engineering made using a little PySpark cluster (Python). An example of the features of Spark SQL may be found on this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\">EDA</a>.</p>\n\n<p><em><strong>User profile</strong></em>\n(based on page_views)</p>\n\n<ul>\n<li>User has previously viewed the ad document? (page_views)</li>\n<li>User page views count</li>\n<li>Categories, Topics and Entities of the documents users have previously viewed (weighted by confidence and TF-IDF) , to model users preferences in a Content-Based Filtering approach. <br>\nPs. According to this <a href=\"https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\">EDA Kernel</a>, about 75% of users in events dataset had at least one additional view on page views dataset (besides the clicks logged as events)</li>\n</ul>\n\n<p><em><strong>Documents and Ads</strong></em></p>\n\n<ul>\n<li>Days elapsed since Ad document and Event (landing page) document were published</li>\n<li>Avg page views by distinct users</li>\n<li>Documents and Ads views count</li>\n</ul>\n\n<p><em><strong>Events</strong></em></p>\n\n<ul>\n<li>Event_hour, adjusted for local timezone (based on state geolocation) and binned (morning, afternoon, ...)</li>\n<li>Is_weekend?</li>\n</ul>\n\n<p><em><strong>Categorical fields</strong></em>\nOne-Hot encoding of categorical fields (about resulting 126K features)</p>\n\n<ul>\n<li>ad_id   </li>\n<li>doc_event_id</li>\n<li>doc_ad_id</li>\n<li>ad_advertiser</li>\n<li>doc_ad_category_id</li>\n<li>doc_ad_entity_id</li>\n<li>doc_ad_publisher_id</li>\n<li>doc_ad_source_id</li>\n<li>doc_ad_topic_id</li>\n<li>doc_event_category_id</li>\n<li>doc_event_entity_id</li>\n<li>doc_event_publisher_id</li>\n<li>doc_event_source_id</li>\n<li>doc_event_topic_id</li>\n<li>event_country</li>\n<li>event_country_state</li>\n<li>event_geo_location</li>\n<li>event_platform</li>\n<li>traffic_source</li>\n</ul>\n\n<p><em><strong>Avg CTR</strong></em> <br>\nAverage CTR (#clicks / #views)  based on categorical fields combinations </p>\n\n<ul>\n<li>ad_id  </li>\n<li>document_id  </li>\n<li>publisher_id  </li>\n<li>advertiser_id  </li>\n<li>campain_id  </li>\n<li>doc_event + doc_ad  </li>\n<li>source_id  </li>\n<li>source_id + country  </li>\n<li>entity_id  </li>\n<li>entity_id + country  </li>\n<li>topic_id  </li>\n<li>topic_id + country  </li>\n<li>category_id  </li>\n<li>category_id + country  </li>\n</ul>\n\n<p><em><strong>Content-Based Similarities</strong></em> <br>\nCosine similarity between user profile and ad doc aspects vectors (TF-IDF)  </p>\n\n<ul>\n<li>user_doc_ad_sim_categories  </li>\n<li>user_doc_ad_sim_topics  </li>\n<li>user_doc_ad_sim_entities  </li>\n</ul>\n\n<p>Cosine similarity between event doc (landing page) and ad doc aspects vectors (TF-IDF)</p>\n\n<ul>\n<li>doc_event_doc_ad_sim_categories</li>\n<li>doc_event_doc_ad_sim_topics</li>\n<li>doc_event_doc_ad_sim_entities</li>\n</ul>\n\n<h3>Cross-validation</h3>\n\n<p>I used a fixed validation set with the same days distribution of the test set (20% of clicks of training events in the first 11 days and all events in days 12 and 13). The alignment of my CV and public LB was accurate at the 4th decimal digit. </p>\n\n<h3>1st level models</h3>\n\n<ul>\n<li>LibFFM (hashed categorical fields) - 0.67932</li>\n<li>LibFFM (hashed categorical fields and binned numeric features) - 0.6784</li>\n<li>LibFFM (hashed categorical fields and binned numeric features trained with the latter 30% events (last days) - 0.6736</li>\n<li>FTRL (hashed categorical quadratic interactions) - 0.67659</li>\n<li>VW (-ftrl, hashed quad interactions among categorical + numeric binned) - 0.6751</li>\n<li>VW (-ftrl, hashed quad interactions only among categorical + numeric raw (not binned)) - 0.67691</li>\n<li>LightGBM (only numeric features) - 0.6707</li>\n<li>XGBoost (numeric and OHE categories) - 0.6689</li>\n<li>RankLib (Random Forests, LambdaMART, ListNet, AdaBoost, RankBoost, Coordinate Ascent) (only numeric fieds)  ~ 0.65-0.66 each model</li>\n</ul>\n\n<h3>2nd level (emsembling)</h3>\n\n<p>I am newbie in emsembling and had little time to explore it. <br>\nMy first approach was to try many kinds of weighted averages (arithmetic, geometric, harmonic) of model predictions. The best setting was a weighted average of inverse logarithmic from the top 4 models (LB=0.68418).</p>\n\n<p>My next approach was to apply XGBoost with rank objective, which is a natural choice as it optimizes the contest metric (MAP), to emsemble models predictions and some numeric features. \nAs I worked with a fixed validation set in 1st level, I had to train model only using validation set (blending) model predictions and original features.  So, I split validation set in train and eval set (50% each), where the eval also followed the test set days distribution. I excluded leaked rows for training, to let models learn other aspects as leak was free :). Using my 6 top models predictions and 15 numeric features I got my best LB score yesterday: <em>0.68716</em>.</p>\n\n<p>I also trained 9 FFM models on subsets of the most frequent event countries (US, CA, GB, AU, Other) and states (CA, FL, TX, NY), and added those predictions and equivalente OHE geo categories in the emsemble, but could not finish before the competition end. I will try an emsemble with all my models in the next few days to see how it would have been.</p>\n\n<p>Kaggle is really a university for cutting edge Machine Learning.</p>\n\n<h3>What didn't work</h3>\n\n<ul>\n<li>Spark ALS Matrix Factorization - User-Based Collaborative Filtering with rank=10 (the maximum I could use in a Spark cluster) underperformed (MAP=0.56). Probably because of sparseness of data.</li>\n<li>One-Hot Encoding Categorical Features - I used OHE as input for XGBoost, but according to better performing models of other competitors using raw id of categorical fields would be fine for trees emsembles. OHE are costly to build, because they require lots of dictionaries. And FFM, FTRL and VW worked well with hashed features.</li>\n<li>RankLib is a collection of algorithms for ranking, using the same input format, which in nice. Therefore, their Java implementation use lots of memory, so I could run on just a sample o numeric data. And their predictions underperformed compared to XGBoost and LightGBM.</li>\n<li>Using GBDT leaf nodes as features to FFM (winning approach in Criteo competition, by 3 idiots) did not worked for me. They actually reduced by CV score.</li>\n<li>Keeping a fixed validation set using the same test set days distribution allowed a CV very aligned with LB. Therefore, in the 2nd level (emsemble), I had to train model using only validation set (blending) model predictions and engineered features. This blending was my best approach, but studying in the forums, I saw that it could be better to run out-of-fold predictions for fixed folds in the 1st level. This way, I would be able to use the full training set, stacking in the 2nd level model predictions (without leak) together with engineered features.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157397,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "01/20/2017 19:20:12",
      "content": "<p>I made available my part of our solution at <a href=\"https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle\">https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle</a>. The README file describes the approach in more details.</p>\n\n<p>In short, my contribution was 5 models on the first level: SVM, FTRL, XGB and ET on mean target value features, FFM on XGB leaves. There were some other models like VW or FFM without leaves, but at the end they didn't contribute much to the ensemble, so they aren't included in the code. \nAlso we had 3 more models from Diaman (I'll let him describe his approach).</p>\n\n<p>The second level model was an XGB with pairwise loss, which we trained only on the half of the data. </p>\n\n<p>Before merging we were around ~40 position, but when we combined the models, we jumped to 15th. After that going up from 15th was really tough. </p>\n\n<p>I learned a lot from this competition, thanks everybody. Also, big thanks to my teammate - it was a lot of fun. </p>\n\n<p>See you all in the next competitions ;-) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "157058": "Congratulations code monkey, brain-afk and Three Data Points for the top 3 finish. Congratulations to other top finishers as well. \r\n\r\nIt was a very interesting competition due to lot of factors such as Data size, number of tables to use, test set having both in time and out time samples, ensembling for map metric etc\r\n\r\nIt will be really helpful for other people in the community if the top finishers share their ideas, codes and the thinking that went behind them in this thread.\r\n\r\nThank you Kaggle and outbrain for this nice competition. Thanks to my team mates and everyone else for making this competition more lively.\r\n\r\nThank you.!",
    "157062": "Big congrats to the top 2 finishes!! Both did amazing! And also huge congrats to Andrii Cherednychenko for getting the score playing solo.  I was actually thinking about merging with you on the last day of merging deadline.\r\n\r\nCarl and CuteChibiko should actually be given all the credits to. I just did some trivial stuff and tried and failed a bunch of ideas.\r\n\r\nSo I will briefly describe our approaches and let Carl and CuteChibiko dive into details later.\r\n\r\nour final approach in one sentence:\r\nffm models + xgboost models as first layer, and xgboost as second layer\r\n\r\nfor ffm models, we have two loss functions, softmax loss (that distributed within each group) and pairwise rank loss.\r\nfor xgboost models, it is just pairwise rank.\r\n\r\nI was focusing on ftrl models for a while at the beginning of the competition. So I tried lambda rank and pairwise rank with them. If I select two way interactions carefully, the ftrl model performance is actually not too far away from ffm model but it can take more than 10+ hours to generate and it contributed little to second layer ensemble, so I moved on.\r\n\r\nIn the last week, I was focusing on nn models. It gets worse result than ffm model but about the same as xgboost model. However one single run of nn models doesn't really give a very stable result (even with 50m data points!!), and I didn't really have time for generating say, 10 - 20 nn models and then average them, so I gave up there too.",
    "157107": "i also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. \r\n\r\nfor the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. \r\n\r\ni also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.",
    "157115": "Congrats to the winners, and well done to all participants. :-)\r\n\r\nMy solution is a simple average of 3 diversified xgboost models. \r\n\r\n**Primary features**\r\n\r\n- Calculation of click ratio for categorical features\r\n- Count of ads per display_id\r\n- Count of page_views for ad-related categoricals\r\n- Bucketing of numerical features\r\n\r\n\r\nBest xgboost\r\nPB: 0.68461\r\n\r\nThe models contained between 40-60 features and were trained on 90% data, with the remaining 10% as validation.\r\n\r\n**Best models parameters**\r\n\r\n- Metric: logloss\r\n- Colsample: 0.30\r\n- Subsample: 0.85\r\n- Depth: 7\r\n- ETA: 0.1 (or less, can't remember)\r\n- Child weight: 5\r\n- Gamma: 0\r\n- Rounds: ~5.000",
    "157119": "[quote=Sameh Faidi;157107]\r\n\r\ni also used xgboost with pairwise and map@12 metric. but couldn't go over 0.686 with a single model. i guess it comes down to feature engineering. \r\n\r\nfor the page views i extracted 3 features that helped a lot. 1) if the user viewed the ad landing page. 2) if the user viewed any page from the same source as the ad landing page 3) if the user viewed any page from the same publisher as the ad landing page. \r\n\r\ni also created some features based on the time when the user viewed the landing page but for some reason this didnt improve the score.\r\n\r\n[/quote]\r\n\r\nCould you explain feature (1)? By ad landing page do you mean the page where the context is displayed? If I remember correctly, for almost all contexts one could find such user-page record in page views, is that what you saw as well?\r\n\r\nThe time features I created didn't help with my fm models, but was useful in my second level xgboost model.",
    "157168": "Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?",
    "157182": "[quote=FengLi;157168]\r\n\r\n@Frederik Cool! You set 5000 rounds? How did you handle the category , topic and entity  variable with the confidence level. For CTR feature, how did you handle NAN?\r\n\r\n[/quote]\r\n\r\nThe approximately 5.000 rounds were determined by auto stopping. (10% data used for validation)\r\n\r\nAmong the categorical variables with multiple values, I selected the single variable with the highest confidence level, and added the confidence level as a bucketed variable. Any NaN's are set to \"-1\" (out of range).",
    "157204": "5000 rounds is impressive! How long does it take to train your model and what resources did you use?",
    "157215": "This Kaggle competition was special for me, as the first I was really engaged on. Three months with lots of learning and little sleep. To be in the first LB page scroll (19th) was rewarding for me.\nHere is my detailed solution (questions and feedbacks are welcome):\n\nPs. More details about my approach in this [post series][1].\n\n### Feature Engineering \n\nAll feature engineering made using a little PySpark cluster (Python). An example of the features of Spark SQL may be found on this [EDA][2].\n\n***User profile***\n(based on page_views)\n\n - User has previously viewed the ad document? (page_views)\n - User page views count\n - Categories, Topics and Entities of the documents users have previously viewed (weighted by confidence and TF-IDF) , to model users preferences in a Content-Based Filtering approach.  \nPs. According to this [EDA Kernel][3], about 75% of users in events dataset had at least one additional view on page views dataset (besides the clicks logged as events)\n\n***Documents and Ads***\n\n - Days elapsed since Ad document and Event (landing page) document were published\n - Avg page views by distinct users\n - Documents and Ads views count\n\n***Events***\n\n - Event_hour, adjusted for local timezone (based on state geolocation) and binned (morning, afternoon, ...)\n - Is_weekend?\n\n***Categorical fields***\nOne-Hot encoding of categorical fields (about resulting 126K features)\n\n - ad_id   \n - doc_event_id\n - doc_ad_id\n - ad_advertiser\n - doc_ad_category_id\n - doc_ad_entity_id\n - doc_ad_publisher_id\n - doc_ad_source_id\n - doc_ad_topic_id\n - doc_event_category_id\n - doc_event_entity_id\n - doc_event_publisher_id\n - doc_event_source_id\n - doc_event_topic_id\n - event_country\n - event_country_state\n - event_geo_location\n - event_platform\n - traffic_source\n\n***Avg CTR***  \nAverage CTR (\\#clicks / \\#views)  based on categorical fields combinations \n \n - ad_id  \n - document_id  \n - publisher_id  \n - advertiser_id  \n - campain_id  \n - doc_event + doc_ad  \n - source_id  \n - source_id + country  \n - entity_id  \n - entity_id + country  \n - topic_id  \n - topic_id + country  \n - category_id  \n - category_id + country  \n\n***Content-Based Similarities***  \nCosine similarity between user profile and ad doc aspects vectors (TF-IDF)  \n\n - user_doc_ad_sim_categories  \n - user_doc_ad_sim_topics  \n - user_doc_ad_sim_entities  \n\nCosine similarity between event doc (landing page) and ad doc aspects vectors (TF-IDF)\n\n - doc_event_doc_ad_sim_categories\n - doc_event_doc_ad_sim_topics\n - doc_event_doc_ad_sim_entities\n\n### Cross-validation  \n\nI used a fixed validation set with the same days distribution of the test set (20% of clicks of training events in the first 11 days and all events in days 12 and 13). The alignment of my CV and public LB was accurate at the 4th decimal digit. \n\n### 1st level models \n\n - LibFFM (hashed categorical fields) - 0.67932\n - LibFFM (hashed categorical fields and binned numeric features) - 0.6784\n - LibFFM (hashed categorical fields and binned numeric features trained with the latter 30% events (last days) - 0.6736\n - FTRL (hashed categorical quadratic interactions) - 0.67659\n - VW (-ftrl, hashed quad interactions among categorical + numeric binned) - 0.6751\n - VW (-ftrl, hashed quad interactions only among categorical + numeric raw (not binned)) - 0.67691\n - LightGBM (only numeric features) - 0.6707\n - XGBoost (numeric and OHE categories) - 0.6689\n - RankLib (Random Forests, LambdaMART, ListNet, AdaBoost, RankBoost, Coordinate Ascent) (only numeric fieds)  ~ 0.65-0.66 each model\n\n### 2nd level (emsembling)\nI am newbie in emsembling and had little time to explore it.  \nMy first approach was to try many kinds of weighted averages (arithmetic, geometric, harmonic) of model predictions. The best setting was a weighted average of inverse logarithmic from the top 4 models (LB=0.68418).\n\nMy next approach was to apply XGBoost with rank objective, which is a natural choice as it optimizes the contest metric (MAP), to emsemble models predictions and some numeric features. \nAs I worked with a fixed validation set in 1st level, I had to train model only using validation set (blending) model predictions and original features.  So, I split validation set in train and eval set (50% each), where the eval also followed the test set days distribution. I excluded leaked rows for training, to let models learn other aspects as leak was free :). Using my 6 top models predictions and 15 numeric features I got my best LB score yesterday: *0.68716*.\n\nI also trained 9 FFM models on subsets of the most frequent event countries (US, CA, GB, AU, Other) and states (CA, FL, TX, NY), and added those predictions and equivalente OHE geo categories in the emsemble, but could not finish before the competition end. I will try an emsemble with all my models in the next few days to see how it would have been.\n\nKaggle is really a university for cutting edge Machine Learning.\n\n### What didn't work\n- Spark ALS Matrix Factorization - User-Based Collaborative Filtering with rank=10 (the maximum I could use in a Spark cluster) underperformed (MAP=0.56). Probably because of sparseness of data.\n- One-Hot Encoding Categorical Features - I used OHE as input for XGBoost, but according to better performing models of other competitors using raw id of categorical fields would be fine for trees emsembles. OHE are costly to build, because they require lots of dictionaries. And FFM, FTRL and VW worked well with hashed features.\n- RankLib is a collection of algorithms for ranking, using the same input format, which in nice. Therefore, their Java implementation use lots of memory, so I could run on just a sample o numeric data. And their predictions underperformed compared to XGBoost and LightGBM.\n- Using GBDT leaf nodes as features to FFM (winning approach in Criteo competition, by 3 idiots) did not worked for me. They actually reduced by CV score.\n- Keeping a fixed validation set using the same test set days distribution allowed a CV very aligned with LB. Therefore, in the 2nd level (emsemble), I had to train model using only validation set (blending) model predictions and engineered features. This blending was my best approach, but studying in the forums, I saw that it could be better to run out-of-fold predictions for fixed folds in the 1st level. This way, I would be able to use the full training set, stacking in the 2nd level model predictions (without leak) together with engineered features.\n\n\n  [1]: https://medium.com/unstructured/how-feature-engineering-can-help-you-do-well-in-a-kaggle-competition-part-i-9cc9a883514d\n  [2]: https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion\n  [3]: https://www.kaggle.com/gspmoreira/outbrain-click-prediction/unveiling-page-views-csv-with-pyspark/discussion",
    "157322": "[quote=sswt;157204]\r\n\r\n5000 rounds is impressive! How long does it take to train your model and what resources did you use?\r\n\r\n[/quote]\r\n\r\nIt took 4-5 days on a 6 core CPU.",
    "157397": "I made available my part of our solution at [https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle][1]. The README file describes the approach in more details.\r\n\r\nIn short, my contribution was 5 models on the first level: SVM, FTRL, XGB and ET on mean target value features, FFM on XGB leaves. There were some other models like VW or FFM without leaves, but at the end they didn't contribute much to the ensemble, so they aren't included in the code. \r\nAlso we had 3 more models from Diaman (I'll let him describe his approach).\r\n\r\nThe second level model was an XGB with pairwise loss, which we trained only on the half of the data. \r\n\r\nBefore merging we were around ~40 position, but when we combined the models, we jumped to 15th. After that going up from 15th was really tough. \r\n\r\nI learned a lot from this competition, thanks everybody. Also, big thanks to my teammate - it was a lot of fun. \r\n\r\nSee you all in the next competitions ;-) \r\n\r\n  [1]: https://github.com/alexeygrigorev/outbrain-click-prediction-kaggle"
  },
  "source": "meta"
}