{
  "id": 25399,
  "title": "Best single model performance",
  "url": "/competitions/outbrain-click-prediction/discussion/25399",
  "author_name": "SRK",
  "post_date": "2016-11-13T16:39:01.087000",
  "votes": 10,
  "comment_count": 55,
  "views": 6108,
  "content": "<p>A classic thread.!</p>\n\n<p>To start with, our current LB score (0.67455) is from a single FTRL model.</p>",
  "messages": [
    {
      "id": 144296,
      "postDate": "2016-11-13T16:39:01.087Z",
      "content": "<p>A classic thread.!</p>\n\n<p>To start with, our current LB score (0.67455) is from a single FTRL model.</p>",
      "rawMarkdown": "A classic thread.!\r\n\r\nTo start with, our current LB score (0.67455) is from a single FTRL model.",
      "votes": 10
    },
    {
      "id": 157040,
      "postDate": "2017-01-19T00:32:34.027Z",
      "content": "<p>So, we've just checked our best single and it got 0.70017 Public / 0.70053 Private. </p>\n\n<p>It's a 2 times bagged on-disk FFM (my own implementation), which was trained for several hours and used less then 8Gb of memory.</p>",
      "rawMarkdown": "So, we've just checked our best single and it got 0.70017 Public / 0.70053 Private. \r\n\r\nIt's a 2 times bagged on-disk FFM (my own implementation), which was trained for several hours and used less then 8Gb of memory.",
      "votes": 8
    },
    {
      "id": 144425,
      "postDate": "2016-11-14T08:06:12.677Z",
      "content": "<p>FFM 0.689 using 24 variables.</p>",
      "rawMarkdown": "FFM 0.689 using 24 variables.",
      "votes": 5
    },
    {
      "id": 157185,
      "postDate": "2017-01-19T15:31:26.640Z",
      "content": "<p>@Sangxia</p>\n\n<p>Objective was just logloss.</p>\n\n<p>Features - a lot:</p>\n\n<ul>\n<li>All what can be joined to clicks from other tables - document, its categories, topics, source, publisher, same for ad doc, platform, location, campaign, advertiser, uid, weekday, hour</li>\n<li>Different counters on prev/future user behaviour - did he viewed this ad, ad with the same doc, ad with doc from the same source, and etc, did he clicked it, same for future</li>\n<li>Information from page views - leak, documents user viewed, document sources user viewed, documents he viewed in one hour after click</li>\n<li>Some special features like diff between ad document creation time and view time, diff between ad document creation time and view document creation time</li>\n</ul>\n\n<p>Maybe I missed something for now, but we'll write more thorough post with our solution later and share the code.</p>",
      "rawMarkdown": "@Sangxia\r\n\r\nObjective was just logloss.\r\n\r\nFeatures - a lot:\r\n\r\n* All what can be joined to clicks from other tables - document, its categories, topics, source, publisher, same for ad doc, platform, location, campaign, advertiser, uid, weekday, hour\r\n* Different counters on prev/future user behaviour - did he viewed this ad, ad with the same doc, ad with doc from the same source, and etc, did he clicked it, same for future\r\n* Information from page views - leak, documents user viewed, document sources user viewed, documents he viewed in one hour after click\r\n* Some special features like diff between ad document creation time and view time, diff between ad document creation time and view document creation time\r\n\r\nMaybe I missed something for now, but we'll write more thorough post with our solution later and share the code.",
      "votes": 3
    },
    {
      "id": 157102,
      "postDate": "2017-01-19T07:48:16.610Z",
      "content": "<p>I was pretty much on single ffm model, just averaged last 5 computations with slightly different features. Devil is in details though, will write about my approach in solutions sharing thread</p>",
      "rawMarkdown": "I was pretty much on single ffm model, just averaged last 5 computations with slightly different features. Devil is in details though, will write about my approach in solutions sharing thread",
      "votes": 3
    },
    {
      "id": 156248,
      "postDate": "2017-01-15T07:26:18.240Z",
      "content": "<p>@Benjamin Chu: yes i am using rank pairwise objective and map@12 as eval metric. just remember to set the group when contrusting the dmatrix. You can also get similar performance using simple binary classification. my best single xgb is now scoreing about 0.684 in LB (local cv score 0.688).  i am not doing anything special with categorical features as most of them can be left as they are (numerical presentation). </p>",
      "rawMarkdown": "@Benjamin Chu: yes i am using rank pairwise objective and map@12 as eval metric. just remember to set the group when contrusting the dmatrix. You can also get similar performance using simple binary classification. my best single xgb is now scoreing about 0.684 in LB (local cv score 0.688).  i am not doing anything special with categorical features as most of them can be left as they are (numerical presentation). ",
      "votes": 3
    },
    {
      "id": 155129,
      "postDate": "2017-01-09T18:23:21.400Z",
      "content": "<p>My best single models scores on LB:</p>\n\n<p>FFM with 21 categorical features + leak                               = 0.67932\nLightGBM (lambdarank) with 74 numeric features + leak = 0.6707</p>\n\n<p>Trying to emsemble LightGBM, and some weaker models (XGBoost and RankLib) with FFM many ways (many types scores and ranks averages, Logistic Regression, GBDT leafs index and score bins), without being able to beat my single FFM model up to now.</p>",
      "rawMarkdown": "My best single models scores on LB:\r\n\r\nFFM with 21 categorical features + leak                               = 0.67932\r\nLightGBM (lambdarank) with 74 numeric features + leak = 0.6707\r\n\r\nTrying to emsemble LightGBM, and some weaker models (XGBoost and RankLib) with FFM many ways (many types scores and ranks averages, Logistic Regression, GBDT leafs index and score bins), without being able to beat my single FFM model up to now.",
      "votes": 3
    },
    {
      "id": 157101,
      "postDate": "2017-01-19T07:43:30.947Z",
      "content": "<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>",
      "rawMarkdown": "single [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM",
      "votes": 4,
      "replies": [
        {
          "id": 157283,
          "postDate": "2017-01-20T02:57:57.013Z",
          "content": "<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]</p>\n\n<p>hi, i tried gbm a lot ... i use origin features about 18 without ad_id,display_id ,got 0.65...and then  i add counts feature of cate_uuid, topic_uuid , entity_uuid( i transform them to one later), and click_rate of cate , topic , entity, but the map@12 decrease greatly ... i cannot figure out why\ncould u share your solution for learing ? </p>",
          "rawMarkdown": "[quote=insulator;157101]\r\n\r\nsingle [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM\r\n\r\n[/quote]\r\n\r\nhi, i tried gbm a lot ... i use origin features about 18 without ad_id,display_id ,got 0.65...and then  i add counts feature of cate_uuid, topic_uuid , entity_uuid( i transform them to one later), and click_rate of cate , topic , entity, but the map@12 decrease greatly ... i cannot figure out why\r\ncould u share your solution for learing ? "
        },
        {
          "id": 157284,
          "postDate": "2017-01-20T03:07:15.450Z",
          "content": "<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>",
          "rawMarkdown": "[quote=insulator;157101]\r\n\r\nsingle [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM\r\n\r\n[/quote]\r\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?\r\n\r\n\r\nThank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.",
          "replies": [
            {
              "id": 157298,
              "postDate": "2017-01-20T05:47:20.767Z",
              "content": "<p>[quote=Little Boat;157284]</p>\n\n<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>\n\n<p>[/quote]</p>\n\n<ol>\n<li>partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. </li>\n<li>features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.</li>\n<li>custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. </li>\n<li>not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. </li>\n</ol>\n\n<p>I think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). </p>\n\n<p>But maybe LightGBM is not so easy to tune, especially \"num_leaves\". </p>",
              "rawMarkdown": "[quote=Little Boat;157284]\r\n\r\n[quote=insulator;157101]\r\n\r\nsingle [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM\r\n\r\n[/quote]\r\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?\r\n\r\n\r\nThank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.\r\n\r\n[/quote]\r\n\r\n1. partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. \r\n2. features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.\r\n3. custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. \r\n4. not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. \r\n\r\nI think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). \r\n\r\nBut maybe LightGBM is not so easy to tune, especially \"num_leaves\". \r\n",
              "votes": 2
            },
            {
              "id": 157300,
              "postDate": "2017-01-20T06:17:18.570Z",
              "content": "<p>[quote=insulator;157298]</p>\n\n<p>[quote=Little Boat;157284]</p>\n\n<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>\n\n<p>[/quote]</p>\n\n<ol>\n<li>partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. </li>\n<li>features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.</li>\n<li>custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. </li>\n<li>not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. </li>\n</ol>\n\n<p>I think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). </p>\n\n<p>But maybe LightGBM is not so easy to tune, especially \"num_leaves\". </p>\n\n<p>[/quote]\nThank you very much!!!\nSo with the same feature set, I can easily get 0.68+ with xgboost pairwise rank without tuning but lightgbm only gave about 0.65 ish. I didn't modify the objective function but just wrote a script that saves the model result and evaluates map@12 every 20 rounds, and stops training when map@12 stops improving.</p>\n\n<p>I can think of two reasons that I got much worse scores than you</p>\n\n<p>1)  usually we do feature engineering and only include features that improve the score. I was adding features to xgboost so the final feature set are kind of \"optimized\" for xgboost, but it became much worse with lightgbm</p>\n\n<p>2)  your customized objective function is much better than lambdarank for map@12 or this specific dataset</p>\n\n<p>I think I tried num_leaves = 256 and 128 with other parameters being default, so it shouldn't have been parameter tuning?</p>\n\n<p>It would be great if you can give one or two tips on how to \"better use\" lightgbm :)</p>",
              "rawMarkdown": "[quote=insulator;157298]\r\n\r\n[quote=Little Boat;157284]\r\n\r\n[quote=insulator;157101]\r\n\r\nsingle [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM\r\n\r\n[/quote]\r\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?\r\n\r\n\r\nThank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.\r\n\r\n[/quote]\r\n\r\n1. partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. \r\n2. features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.\r\n3. custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. \r\n4. not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. \r\n\r\nI think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). \r\n\r\nBut maybe LightGBM is not so easy to tune, especially \"num_leaves\". \r\n\r\n\r\n[/quote]\r\nThank you very much!!!\r\nSo with the same feature set, I can easily get 0.68+ with xgboost pairwise rank without tuning but lightgbm only gave about 0.65 ish. I didn't modify the objective function but just wrote a script that saves the model result and evaluates map@12 every 20 rounds, and stops training when map@12 stops improving.\r\n\r\nI can think of two reasons that I got much worse scores than you\r\n\r\n1)  usually we do feature engineering and only include features that improve the score. I was adding features to xgboost so the final feature set are kind of \"optimized\" for xgboost, but it became much worse with lightgbm\r\n\r\n2)  your customized objective function is much better than lambdarank for map@12 or this specific dataset\r\n\r\nI think I tried num_leaves = 256 and 128 with other parameters being default, so it shouldn't have been parameter tuning?\r\n\r\nIt would be great if you can give one or two tips on how to \"better use\" lightgbm :)"
            },
            {
              "id": 157302,
              "postDate": "2017-01-20T06:29:44.823Z",
              "content": "<p>[quote=Little Boat;157300]</p>\n\n<p>[quote=insulator;157298]</p>\n\n<p>[quote=Little Boat;157284]</p>\n\n<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>\n\n<p>[/quote]</p>\n\n<ol>\n<li>partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. </li>\n<li>features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.</li>\n<li>custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. </li>\n<li>not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. </li>\n</ol>\n\n<p>I think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). </p>\n\n<p>But maybe LightGBM is not so easy to tune, especially \"num_leaves\". </p>\n\n<p>[/quote]\nThank you very much!!!\nSo with the same feature set, I can easily get 0.68+ with xgboost pairwise rank without tuning but lightgbm only gave about 0.65 ish. I didn't modify the objective function but just wrote a script that saves the model result and evaluates map@12 every 20 rounds, and stops training when map@12 stops improving.</p>\n\n<p>I can think of two reasons that I got much worse scores than you</p>\n\n<p>1)  usually we do feature engineering and only include features that improve the score. I was adding features to xgboost so the final feature set are kind of \"optimized\" for xgboost, but it became much worse with lightgbm</p>\n\n<p>2)  your customized objective function is much better than lambdarank for map@12 or this specific dataset</p>\n\n<p>I think I tried num_leaves = 256 and 128 with other parameters being default, so it shouldn't have been parameter tuning?</p>\n\n<p>It would be great if you can give one or two tips on how to \"better use\" lightgbm :)</p>\n\n<p>[/quote]</p>\n\n<p>Actually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)</p>\n\n<p>I think the most useful parameters are : <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters</a></p>\n\n<p>e.g. limit the min data in one leaves and so on.</p>\n\n<p>And I think another reason is over-fitting. \nLightGBM is more easy to cause over-fitting due to it fits training data better. \nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  </p>",
              "rawMarkdown": "[quote=Little Boat;157300]\r\n\r\n[quote=insulator;157298]\r\n\r\n[quote=Little Boat;157284]\r\n\r\n[quote=insulator;157101]\r\n\r\nsingle [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM\r\n\r\n[/quote]\r\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?\r\n\r\n\r\nThank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.\r\n\r\n[/quote]\r\n\r\n1. partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. \r\n2. features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.\r\n3. custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. \r\n4. not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. \r\n\r\nI think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). \r\n\r\nBut maybe LightGBM is not so easy to tune, especially \"num_leaves\". \r\n\r\n\r\n[/quote]\r\nThank you very much!!!\r\nSo with the same feature set, I can easily get 0.68+ with xgboost pairwise rank without tuning but lightgbm only gave about 0.65 ish. I didn't modify the objective function but just wrote a script that saves the model result and evaluates map@12 every 20 rounds, and stops training when map@12 stops improving.\r\n\r\nI can think of two reasons that I got much worse scores than you\r\n\r\n1)  usually we do feature engineering and only include features that improve the score. I was adding features to xgboost so the final feature set are kind of \"optimized\" for xgboost, but it became much worse with lightgbm\r\n\r\n2)  your customized objective function is much better than lambdarank for map@12 or this specific dataset\r\n\r\nI think I tried num_leaves = 256 and 128 with other parameters being default, so it shouldn't have been parameter tuning?\r\n\r\nIt would be great if you can give one or two tips on how to \"better use\" lightgbm :)\r\n\r\n[/quote]\r\n\r\nActually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)\r\n\r\nI think the most useful parameters are : https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\r\n\r\ne.g. limit the min data in one leaves and so on.\r\n\r\nAnd I think another reason is over-fitting. \r\nLightGBM is more easy to cause over-fitting due to it fits training data better. \r\nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \r\nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  \r\n",
              "votes": 2
            },
            {
              "id": 157303,
              "postDate": "2017-01-20T06:34:38.857Z",
              "content": "<p>[quote=insulator;157302]</p>\n\n<p>Actually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)</p>\n\n<p>I think the most useful parameters are : <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters</a></p>\n\n<p>e.g. limit the min data in one leaves and so on.</p>\n\n<p>And I think another reason is over-fitting. \nLightGBM is more easy to cause over-fitting due to it fits training data better. \nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  </p>\n\n<p>[/quote]</p>\n\n<p>That is a good idea! Thanks, will definitely try next time. Will MAP become one of the default evaluation options for lightgbm in future?</p>",
              "rawMarkdown": "[quote=insulator;157302]\r\n\r\nActually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)\r\n\r\nI think the most useful parameters are : https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\r\n\r\ne.g. limit the min data in one leaves and so on.\r\n\r\nAnd I think another reason is over-fitting. \r\nLightGBM is more easy to cause over-fitting due to it fits training data better. \r\nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \r\nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  \r\n\r\n\r\n[/quote]\r\n\r\nThat is a good idea! Thanks, will definitely try next time. Will MAP become one of the default evaluation options for lightgbm in future?"
            },
            {
              "id": 157304,
              "postDate": "2017-01-20T06:39:15.457Z",
              "content": "<p>[quote=Little Boat;157303]</p>\n\n<p>[quote=insulator;157302]</p>\n\n<p>Actually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)</p>\n\n<p>I think the most useful parameters are : <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters</a></p>\n\n<p>e.g. limit the min data in one leaves and so on.</p>\n\n<p>And I think another reason is over-fitting. \nLightGBM is more easy to cause over-fitting due to it fits training data better. \nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  </p>\n\n<p>[/quote]</p>\n\n<p>That is a good idea! Thanks, will definitely try next time. Will MAP become one of the default evaluation options for lightgbm in future?</p>\n\n<p>[/quote]</p>\n\n<p>Sure, </p>\n\n<p>BTW, if you met any parameter tuning issues or need new features,  welcome to open issues on github. </p>",
              "rawMarkdown": "[quote=Little Boat;157303]\r\n\r\n[quote=insulator;157302]\r\n\r\nActually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)\r\n\r\nI think the most useful parameters are : https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\r\n\r\ne.g. limit the min data in one leaves and so on.\r\n\r\nAnd I think another reason is over-fitting. \r\nLightGBM is more easy to cause over-fitting due to it fits training data better. \r\nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \r\nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  \r\n\r\n\r\n[/quote]\r\n\r\nThat is a good idea! Thanks, will definitely try next time. Will MAP become one of the default evaluation options for lightgbm in future?\r\n\r\n[/quote]\r\n\r\nSure, \r\n\r\nBTW, if you met any parameter tuning issues or need new features,  welcome to open issues on github. "
            },
            {
              "id": 157305,
              "postDate": "2017-01-20T06:40:15.023Z",
              "content": "<p>[quote=insulator;157304]</p>\n\n<p>Sure, </p>\n\n<p>BTW, if you met any parameter tuning issues or need new features,  welcome to open issues on github. </p>\n\n<p>[/quote]</p>\n\n<p>Definitely will!!</p>\n\n<p>Thanks again! </p>",
              "rawMarkdown": "[quote=insulator;157304]\r\n\r\nSure, \r\n\r\nBTW, if you met any parameter tuning issues or need new features,  welcome to open issues on github. \r\n\r\n[/quote]\r\n\r\nDefinitely will!!\r\n\r\nThanks again! \r\n"
            }
          ]
        }
      ]
    },
    {
      "id": 154849,
      "postDate": "2017-01-08T10:48:50.273Z",
      "content": "<p>am I the only one using xgboost? :) untuned single model with about 30 features LB 0.677</p>",
      "rawMarkdown": "am I the only one using xgboost? :) untuned single model with about 30 features LB 0.677",
      "votes": 4
    },
    {
      "id": 144947,
      "postDate": "2016-11-16T03:21:51.983Z",
      "content": "<p>BPR: 0.64830, FM using 10 features + leak: 0.67539</p>",
      "rawMarkdown": "BPR: 0.64830, FM using 10 features + leak: 0.67539",
      "votes": 3
    },
    {
      "id": 144731,
      "postDate": "2016-11-15T07:53:23.217Z",
      "content": "<p>0.67928  - my current LB score of single ftrl, ~15 features + some interactions</p>",
      "rawMarkdown": "0.67928  - my current LB score of single ftrl, ~15 features + some interactions",
      "votes": 4
    },
    {
      "id": 157070,
      "postDate": "2017-01-19T03:46:17.737Z",
      "content": "<p>@Alexey Noskov congrats! that's amazing. what features did you use for the single model and what was the objective?</p>",
      "rawMarkdown": "@Alexey Noskov congrats! that's amazing. what features did you use for the single model and what was the objective?",
      "votes": 1
    },
    {
      "id": 155613,
      "postDate": "2017-01-12T01:25:46.463Z",
      "content": "<p>0.674 using single xgboost</p>",
      "rawMarkdown": "0.674 using single xgboost",
      "votes": 1
    },
    {
      "id": 144566,
      "postDate": "2016-11-14T17:05:52.110Z",
      "content": "<p>Thanks Carl. So there is a lot of potential in FFMs then. </p>",
      "rawMarkdown": "Thanks Carl. So there is a lot of potential in FFMs then. ",
      "votes": 1
    },
    {
      "id": 144375,
      "postDate": "2016-11-14T02:45:53.200Z",
      "content": "<p>Gilberto, Great score using just three variables.!</p>\n\n<p>Our best FFM scores 0.669 using multiple variables. </p>\n\n<p>I think people are getting far better scores using FFM than FTRL. This is my first time using FFMs. Any help to improve the FFM scores would be highly appreciated.</p>",
      "rawMarkdown": "Gilberto, Great score using just three variables.!\r\n\r\nOur best FFM scores 0.669 using multiple variables. \r\n\r\nI think people are getting far better scores using FFM than FTRL. This is my first time using FFMs. Any help to improve the FFM scores would be highly appreciated.",
      "votes": 1
    },
    {
      "id": 155142,
      "postDate": "2017-01-09T20:52:44.067Z",
      "content": "<p>I finally managed to complete a run of FFM with leak. Just to confirm everyone's impression, with almost exactly the same features (I had to remove one from FFM), FFM gets a better LB score compared to FTRL.</p>\n\n<ul>\n<li>FTRL: 0.66200</li>\n<li>FFM: 0.66525</li>\n</ul>",
      "rawMarkdown": "I finally managed to complete a run of FFM with leak. Just to confirm everyone's impression, with almost exactly the same features (I had to remove one from FFM), FFM gets a better LB score compared to FTRL.\r\n\r\n- FTRL: 0.66200\r\n- FFM: 0.66525",
      "votes": 2,
      "replies": [
        {
          "id": 155208,
          "postDate": "2017-01-10T04:56:35.360Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 155053,
      "postDate": "2017-01-09T08:28:47.950Z",
      "content": "<p>I am still on single model</p>",
      "rawMarkdown": "I am still on single model",
      "votes": 2
    },
    {
      "id": 154797,
      "postDate": "2017-01-08T04:16:33.827Z",
      "content": "<p>@Guilherme de Oliveira Santos</p>\n\n<p>One of my submission scored 0.6801 is FTRL on 14 features. The FTRL shared in Kernel is awesome but quite slow, so I use <a href=\"https://github.com/jeongyoonlee/Kaggler\">Kaggler</a> package instead (<a href=\"https://github.com/jeongyoonlee/Kaggler#from-source-code\">install from source code</a>). It runs a epoch within 20 mins even when <code>interaction=True</code>. This might help you iterate faster and find a better hyper parameter set.</p>\n\n<p>Edit: Sorry I post the wrong score, 0.6801 is my CV score, its LB is 0.6787</p>",
      "rawMarkdown": "@Guilherme de Oliveira Santos\r\n\r\nOne of my submission scored 0.6801 is FTRL on 14 features. The FTRL shared in Kernel is awesome but quite slow, so I use [Kaggler][1] package instead ([install from source code][2]). It runs a epoch within 20 mins even when `interaction=True`. This might help you iterate faster and find a better hyper parameter set.\r\n\r\nEdit: Sorry I post the wrong score, 0.6801 is my CV score, its LB is 0.6787\r\n\r\n  [1]: https://github.com/jeongyoonlee/Kaggler\r\n  [2]: https://github.com/jeongyoonlee/Kaggler#from-source-code",
      "votes": 2
    },
    {
      "id": 148275,
      "postDate": "2016-12-03T22:50:39.393Z",
      "content": "<p>My results:</p>\n\n<pre><code>Model           CV      LB\nFTRL + Leak     0.67321 0.67291\nFFM             0.66171 0.66135\nFFM + Leak      0.67636 -------\n</code></pre>",
      "rawMarkdown": "My results:\r\n\r\n    Model           CV  \tLB\r\n    FTRL + Leak \t0.67321\t0.67291\r\n    FFM         \t0.66171\t0.66135\r\n    FFM + Leak      0.67636 -------\r\n\r\n ",
      "votes": 2
    },
    {
      "id": 144354,
      "postDate": "2016-11-13T23:44:33.073Z",
      "content": "<p>FFM, LB: 0.66725 using 2 original features + &quot;leak&quot;</p>",
      "rawMarkdown": "FFM, LB: 0.66725 using 2 original features + \"leak\"",
      "votes": 2
    },
    {
      "id": 157205,
      "postDate": "2017-01-19T17:11:26.670Z",
      "content": "<p>@Sangxia we transformed counts with <code>log2(x + 1)</code> transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as <code>log2(x + 1)</code>. But our best ffm was far from 0.69.</p>",
      "rawMarkdown": "@Sangxia we transformed counts with `log2(x + 1)` transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as `log2(x + 1)`. But our best ffm was far from 0.69.",
      "replies": [
        {
          "id": 157210,
          "postDate": "2017-01-19T17:34:12.287Z",
          "content": "<p>[quote=diaman;157205]</p>\n\n<p>@Sangxia we transformed counts with <code>log2(x + 1)</code> transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as <code>log2(x + 1)</code>. But our best ffm was far from 0.69.</p>\n\n<p>[/quote]</p>\n\n<p>I didn't express myself well. What I am seeing in my CV (let's say no leak, only on future data) is that: \nif we only look at rows for users that have less than say 10 rows in page views, the map was about 0.66; \nas this count increases, the map goes up to above 0.7; but then for the users that have more than 70 rows, it rapidly drops to something like 0.63 or lower. Part of my features is a normalized vector of document ids that a user viewed in the page view file. I was wondering if others observed something like that. Hope this makes some sense.</p>",
          "rawMarkdown": "[quote=diaman;157205]\r\n\r\n@Sangxia we transformed counts with `log2(x + 1)` transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as `log2(x + 1)`. But our best ffm was far from 0.69.\r\n\r\n[/quote]\r\n\r\nI didn't express myself well. What I am seeing in my CV (let's say no leak, only on future data) is that: \r\nif we only look at rows for users that have less than say 10 rows in page views, the map was about 0.66; \r\nas this count increases, the map goes up to above 0.7; but then for the users that have more than 70 rows, it rapidly drops to something like 0.63 or lower. Part of my features is a normalized vector of document ids that a user viewed in the page view file. I was wondering if others observed something like that. Hope this makes some sense.",
          "votes": 1
        }
      ]
    },
    {
      "id": 157200,
      "postDate": "2017-01-19T16:42:30.763Z",
      "content": "<p>For teams using FFM, have you looked at your CV map@12 by user page view count? Asking because I noticed mine has a huge drop when the count goes above about 60, and was wondering whether I made a somewhat suboptimal choice of model.</p>",
      "rawMarkdown": "For teams using FFM, have you looked at your CV map@12 by user page view count? Asking because I noticed mine has a huge drop when the count goes above about 60, and was wondering whether I made a somewhat suboptimal choice of model."
    },
    {
      "id": 157198,
      "postDate": "2017-01-19T16:33:30.430Z",
      "content": "<p>@insulator Could you please share the parameters for your LightGBM? Did you include original categorical features or you just transform all of them into CTR? If you could elaborate the features you generate, that'll be great. Thank you.</p>",
      "rawMarkdown": "@insulator Could you please share the parameters for your LightGBM? Did you include original categorical features or you just transform all of them into CTR? If you could elaborate the features you generate, that'll be great. Thank you."
    },
    {
      "id": 155640,
      "postDate": "2017-01-12T06:22:50.830Z",
      "content": "<p>@peixiang, after this competition ends, could you please publish your Xgboost source code?</p>\n\n<p>I didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)</p>",
      "rawMarkdown": "@peixiang, after this competition ends, could you please publish your Xgboost source code?\r\n\r\nI didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)",
      "replies": [
        {
          "id": 155862,
          "postDate": "2017-01-13T02:30:55.273Z",
          "content": "<p>[quote=Yehoshaphat Schellekens;155640]</p>\n\n<p>@peixiang, after this competition ends, could you please publish your Xgboost source code?</p>\n\n<p>I didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)</p>\n\n<p>[/quote]\nActually I did not use one hot encoding to encode categorical data. For example, to represent document's topic, there are 300 topics which means normally I need a vector of 300 to represent one doc's topic. Instead, I just use the \"main\" topic of this document as a feature. Determining the \"main\" topic of each document is kind of based on some simple data analysis and heuristics. Maybe a more accurate \"main\" topic or adding an extra \"main\" topic will improve the score.</p>",
          "rawMarkdown": "[quote=Yehoshaphat Schellekens;155640]\r\n\r\n@peixiang, after this competition ends, could you please publish your Xgboost source code?\r\n\r\nI didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)\r\n\r\n[/quote]\r\nActually I did not use one hot encoding to encode categorical data. For example, to represent document's topic, there are 300 topics which means normally I need a vector of 300 to represent one doc's topic. Instead, I just use the \"main\" topic of this document as a feature. Determining the \"main\" topic of each document is kind of based on some simple data analysis and heuristics. Maybe a more accurate \"main\" topic or adding an extra \"main\" topic will improve the score.",
          "votes": 1
        }
      ]
    },
    {
      "id": 155255,
      "postDate": "2017-01-10T09:53:46.087Z",
      "content": "<p>@infante, FFM stands for Field-aware Factorization Machines. What I am calling FTRL is SRK's kernel (<a href=\"https://www.kaggle.com/sudalairajkumar/outbrain-click-prediction/ftrl-starter-with-leakage-vars\">https://www.kaggle.com/sudalairajkumar/outbrain-click-prediction/ftrl-starter-with-leakage-vars</a>). </p>",
      "rawMarkdown": "@infante, FFM stands for Field-aware Factorization Machines. What I am calling FTRL is SRK's kernel (https://www.kaggle.com/sudalairajkumar/outbrain-click-prediction/ftrl-starter-with-leakage-vars). "
    },
    {
      "id": 155095,
      "postDate": "2017-01-09T15:01:00.320Z",
      "content": "<p>@Grzegorz Sionkowski can you share how many iteration did you do</p>",
      "rawMarkdown": "@Grzegorz Sionkowski can you share how many iteration did you do"
    },
    {
      "id": 155094,
      "postDate": "2017-01-09T14:54:21.610Z",
      "content": "<p>@Grzegorz Sionkowski Could you share your method after comp? I'm very interested in how you do that.</p>",
      "rawMarkdown": "@Grzegorz Sionkowski Could you share your method after comp? I'm very interested in how you do that."
    },
    {
      "id": 155050,
      "postDate": "2017-01-09T08:13:40.967Z",
      "content": "<p>0.650 - single model with 4 original features, no leak, no sophisticated ML method, just simple iterations. </p>\n\n<p>EDIT: 17 iterations</p>",
      "rawMarkdown": "0.650 - single model with 4 original features, no leak, no sophisticated ML method, just simple iterations. \r\n\r\nEDIT: 17 iterations"
    },
    {
      "id": 154922,
      "postDate": "2017-01-08T19:44:19.163Z",
      "content": "<p>@Sameh Faidi You do some transformation to category variables right? </p>",
      "rawMarkdown": "@Sameh Faidi You do some transformation to category variables right? "
    },
    {
      "id": 154874,
      "postDate": "2017-01-08T13:43:18.960Z",
      "content": "<p>@Benjamin Chu Thanks for the reply! I was wondering about model limitations. Right now I'm at LB 0.67222 using 27 features. My model runs 2 epoch in 40 minutes with peak memory usage of 7GB using the FTRL script by tinrtgu.</p>\n\n<p>I was just about to give up the competition my Notebook has only 12GB but I'll try to reach 0.68x now.</p>",
      "rawMarkdown": "@Benjamin Chu Thanks for the reply! I was wondering about model limitations. Right now I'm at LB 0.67222 using 27 features. My model runs 2 epoch in 40 minutes with peak memory usage of 7GB using the FTRL script by tinrtgu.\r\n\r\nI was just about to give up the competition my Notebook has only 12GB but I'll try to reach 0.68x now.",
      "replies": [
        {
          "id": 155207,
          "postDate": "2017-01-10T04:54:08.393Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 154860,
      "postDate": "2017-01-08T12:08:25.767Z",
      "content": "<p>@Sameh Faidi</p>\n\n<p>Do you use pairwise ranking of XGB?</p>",
      "rawMarkdown": "@Sameh Faidi\r\n\r\nDo you use pairwise ranking of XGB?"
    },
    {
      "id": 154851,
      "postDate": "2017-01-08T11:02:48.040Z",
      "content": "<p>@Sameh: You're the first I see using xgboost! I was trying to use it last week but with my limited amount of RAM (16Go), I could not run it my many features so I switched to FFM. And even now I still not able to run it on whole train set with more than 10 features...</p>\n\n<p>You said you had more than 100Go of RAM, what is your maximum usage? Are you running your model on Python or R?</p>",
      "rawMarkdown": "@Sameh: You're the first I see using xgboost! I was trying to use it last week but with my limited amount of RAM (16Go), I could not run it my many features so I switched to FFM. And even now I still not able to run it on whole train set with more than 10 features...\r\n\r\nYou said you had more than 100Go of RAM, what is your maximum usage? Are you running your model on Python or R?\r\n\r\n"
    },
    {
      "id": 154817,
      "postDate": "2017-01-08T06:30:54.577Z",
      "content": "<p>@-puli-\n14 features will have peak memory usage about 21GB. My server's memory is larger than that. I think you may achieve 0.67X with only 5 features, and in that case the data could fit in a 16GB PC.</p>",
      "rawMarkdown": "@-puli-\r\n14 features will have peak memory usage about 21GB. My server's memory is larger than that. I think you may achieve 0.67X with only 5 features, and in that case the data could fit in a 16GB PC."
    },
    {
      "id": 154813,
      "postDate": "2017-01-08T05:34:46.223Z",
      "content": "<p>@Benjamin Chu: That's a nice script. Just wondering ain't loading the entire dataset in libsvm format heavy on computation. What's your memory like.. just curious :)</p>",
      "rawMarkdown": "@Benjamin Chu: That's a nice script. Just wondering ain't loading the entire dataset in libsvm format heavy on computation. What's your memory like.. just curious :)\r\n"
    },
    {
      "id": 154808,
      "postDate": "2017-01-08T05:09:38.903Z",
      "content": "<p>@FengLi \nYep, I think it's the key to winning this competition.</p>",
      "rawMarkdown": "@FengLi \r\nYep, I think it's the key to winning this competition."
    },
    {
      "id": 154799,
      "postDate": "2017-01-08T04:27:12.773Z",
      "content": "<p>@Benjamin Chu Cool!. You must do a lot of feature engineering.</p>",
      "rawMarkdown": "@Benjamin Chu Cool!. You must do a lot of feature engineering."
    },
    {
      "id": 154698,
      "postDate": "2017-01-07T18:14:48.927Z",
      "content": "<p>Does anyone managed to score above 0.68 using FTRL?</p>",
      "rawMarkdown": "Does anyone managed to score above 0.68 using FTRL?"
    },
    {
      "id": 153036,
      "postDate": "2016-12-29T15:40:23.607Z",
      "content": "<p>Hi all,\nBecause of the limitation of the machine, I'm wondering what's the best single model performance without using page_view.csv (except the leak) ? </p>",
      "rawMarkdown": "Hi all,\r\nBecause of the limitation of the machine, I'm wondering what's the best single model performance without using page_view.csv (except the leak) ? \r\n"
    },
    {
      "id": 152466,
      "postDate": "2016-12-26T14:00:19.923Z",
      "content": "<p>@GuoxinLi Please refer the installation instructions and make sure you have all the requirements fulfilled <a href=\"https://github.com/guestwalk/libffm/blob/master/README\">https://github.com/guestwalk/libffm/blob/master/README</a></p>",
      "rawMarkdown": "@GuoxinLi Please refer the installation instructions and make sure you have all the requirements fulfilled https://github.com/guestwalk/libffm/blob/master/README"
    },
    {
      "id": 152465,
      "postDate": "2016-12-26T13:55:31.393Z",
      "content": "<p>@Keerath Jaggi Thank you for your reply, I remove </p>\n\n<p>DFLAG += -DUSEOMP, CXXFLAGS += -fopenmpnow </p>\n\n<p>now I get error</p>\n\n<p>../sdk/graphlab/cppipc/server/comm_server.hpp:22:10: fatal error: \n      'nanosockets/socket_errors.hpp' file not found\ninclude  nanosockets/socket_errors.hpp\n         ^</p>",
      "rawMarkdown": "@Keerath Jaggi Thank you for your reply, I remove \r\n\r\nDFLAG += -DUSEOMP, CXXFLAGS += -fopenmpnow \r\n\r\nnow I get error\r\n\r\n../sdk/graphlab/cppipc/server/comm_server.hpp:22:10: fatal error: \r\n      'nanosockets/socket_errors.hpp' file not found\r\ninclude  nanosockets/socket_errors.hpp\r\n         ^\r\n"
    },
    {
      "id": 152463,
      "postDate": "2016-12-26T13:47:23.160Z",
      "content": "<p>@GuoxinLi\nOpenMP is required by FFM for parallelization.\nIf OpenMP is not available on your\nplatform, then please comment out the following lines in Makefile.</p>\n\n<pre><code>DFLAG += -DUSEOMP\nCXXFLAGS += -fopenmp\n</code></pre>\n\n<p>and then try 'make clean all'</p>",
      "rawMarkdown": "@GuoxinLi\r\nOpenMP is required by FFM for parallelization.\r\nIf OpenMP is not available on your\r\nplatform, then please comment out the following lines in Makefile.\r\n\r\n    DFLAG += -DUSEOMP\r\n    CXXFLAGS += -fopenmp\r\nand then try 'make clean all'"
    },
    {
      "id": 152461,
      "postDate": "2016-12-26T13:43:27.923Z",
      "content": "<p>I have problem to install FFM on my machine, I try to install libffm from <a href=\"https://github.com/turi-code/python-libffm\">https://github.com/turi-code/python-libffm</a>, but it gave me error \" clang: error: unsupported option '-fopenmp' \", does anyone know how to solve it.</p>",
      "rawMarkdown": "I have problem to install FFM on my machine, I try to install libffm from https://github.com/turi-code/python-libffm, but it gave me error \" clang: error: unsupported option '-fopenmp' \", does anyone know how to solve it."
    },
    {
      "id": 155409,
      "postDate": "2017-01-11T03:18:38.743Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 154942,
      "postDate": "2017-01-08T23:36:21.393Z",
      "content": "<p>@Benjamin Chu: Thanks a lot. </p>",
      "rawMarkdown": "@Benjamin Chu: Thanks a lot. "
    }
  ],
  "comments": [
    {
      "id": 157040,
      "author_name": "Alexey Noskov",
      "author_url": "",
      "post_date": "2017-01-19T00:32:34.027000",
      "content": "<p>So, we've just checked our best single and it got 0.70017 Public / 0.70053 Private. </p>\n\n<p>It's a 2 times bagged on-disk FFM (my own implementation), which was trained for several hours and used less then 8Gb of memory.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 144425,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2016-11-14T08:06:12.677000",
      "content": "<p>FFM 0.689 using 24 variables.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 157185,
      "author_name": "Alexey Noskov",
      "author_url": "",
      "post_date": "2017-01-19T15:31:26.640000",
      "content": "<p>@Sangxia</p>\n\n<p>Objective was just logloss.</p>\n\n<p>Features - a lot:</p>\n\n<ul>\n<li>All what can be joined to clicks from other tables - document, its categories, topics, source, publisher, same for ad doc, platform, location, campaign, advertiser, uid, weekday, hour</li>\n<li>Different counters on prev/future user behaviour - did he viewed this ad, ad with the same doc, ad with doc from the same source, and etc, did he clicked it, same for future</li>\n<li>Information from page views - leak, documents user viewed, document sources user viewed, documents he viewed in one hour after click</li>\n<li>Some special features like diff between ad document creation time and view time, diff between ad document creation time and view document creation time</li>\n</ul>\n\n<p>Maybe I missed something for now, but we'll write more thorough post with our solution later and share the code.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 157102,
      "author_name": "Andrii Cherednychenko",
      "author_url": "",
      "post_date": "2017-01-19T07:48:16.610000",
      "content": "<p>I was pretty much on single ffm model, just averaged last 5 computations with slightly different features. Devil is in details though, will write about my approach in solutions sharing thread</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 156248,
      "author_name": "Sameh Faidi",
      "author_url": "",
      "post_date": "2017-01-15T07:26:18.240000",
      "content": "<p>@Benjamin Chu: yes i am using rank pairwise objective and map@12 as eval metric. just remember to set the group when contrusting the dmatrix. You can also get similar performance using simple binary classification. my best single xgb is now scoreing about 0.684 in LB (local cv score 0.688).  i am not doing anything special with categorical features as most of them can be left as they are (numerical presentation). </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 155129,
      "author_name": "Gabriel Moreira",
      "author_url": "",
      "post_date": "2017-01-09T18:23:21.400000",
      "content": "<p>My best single models scores on LB:</p>\n\n<p>FFM with 21 categorical features + leak                               = 0.67932\nLightGBM (lambdarank) with 74 numeric features + leak = 0.6707</p>\n\n<p>Trying to emsemble LightGBM, and some weaker models (XGBoost and RankLib) with FFM many ways (many types scores and ranks averages, Logistic Regression, GBDT leafs index and score bins), without being able to beat my single FFM model up to now.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 157101,
      "author_name": "Guolin Ke",
      "author_url": "",
      "post_date": "2017-01-19T07:43:30.947000",
      "content": "<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 157283,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-01-20T02:57:57.013000",
          "content": "<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]</p>\n\n<p>hi, i tried gbm a lot ... i use origin features about 18 without ad_id,display_id ,got 0.65...and then  i add counts feature of cate_uuid, topic_uuid , entity_uuid( i transform them to one later), and click_rate of cate , topic , entity, but the map@12 decrease greatly ... i cannot figure out why\ncould u share your solution for learing ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 157284,
          "author_name": "Little Boat",
          "author_url": "",
          "post_date": "2017-01-20T03:07:15.450000",
          "content": "<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 157298,
              "author_name": "Guolin Ke",
              "author_url": "",
              "post_date": "2017-01-20T05:47:20.767000",
              "content": "<p>[quote=Little Boat;157284]</p>\n\n<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>\n\n<p>[/quote]</p>\n\n<ol>\n<li>partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. </li>\n<li>features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.</li>\n<li>custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. </li>\n<li>not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. </li>\n</ol>\n\n<p>I think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). </p>\n\n<p>But maybe LightGBM is not so easy to tune, especially \"num_leaves\". </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 157300,
              "author_name": "Little Boat",
              "author_url": "",
              "post_date": "2017-01-20T06:17:18.570000",
              "content": "<p>[quote=insulator;157298]</p>\n\n<p>[quote=Little Boat;157284]</p>\n\n<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>\n\n<p>[/quote]</p>\n\n<ol>\n<li>partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. </li>\n<li>features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.</li>\n<li>custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. </li>\n<li>not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. </li>\n</ol>\n\n<p>I think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). </p>\n\n<p>But maybe LightGBM is not so easy to tune, especially \"num_leaves\". </p>\n\n<p>[/quote]\nThank you very much!!!\nSo with the same feature set, I can easily get 0.68+ with xgboost pairwise rank without tuning but lightgbm only gave about 0.65 ish. I didn't modify the objective function but just wrote a script that saves the model result and evaluates map@12 every 20 rounds, and stops training when map@12 stops improving.</p>\n\n<p>I can think of two reasons that I got much worse scores than you</p>\n\n<p>1)  usually we do feature engineering and only include features that improve the score. I was adding features to xgboost so the final feature set are kind of \"optimized\" for xgboost, but it became much worse with lightgbm</p>\n\n<p>2)  your customized objective function is much better than lambdarank for map@12 or this specific dataset</p>\n\n<p>I think I tried num_leaves = 256 and 128 with other parameters being default, so it shouldn't have been parameter tuning?</p>\n\n<p>It would be great if you can give one or two tips on how to \"better use\" lightgbm :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 157302,
              "author_name": "Guolin Ke",
              "author_url": "",
              "post_date": "2017-01-20T06:29:44.823000",
              "content": "<p>[quote=Little Boat;157300]</p>\n\n<p>[quote=insulator;157298]</p>\n\n<p>[quote=Little Boat;157284]</p>\n\n<p>[quote=insulator;157101]</p>\n\n<p>single <a href=\"https://github.com/Microsoft/LightGBM\">LightGBM</a>  model with some counting and CTR features get 0.69019 in PLB. </p>\n\n<p>[/quote]\nI tried to make lightgbm work but it just didn't work for me. Could you elaborate more on it? Like did you use lambdarank, with what feature sets,  what parameters, etc?</p>\n\n<p>Thank you very much! I really wanna add lightgbm to my toolkit but so far I haven't had much success in several competitions.</p>\n\n<p>[/quote]</p>\n\n<ol>\n<li>partition data, random pick 10M for the training. Rest of them are used for making counting / CTR feature, then join these feature to the training data. </li>\n<li>features are raw ids and ctr features. ctr features are click_cnt, imp_cnt and ctr. And I extract CTR feature for the all ids, include ad, doc and so on. Also extract CTR feature for some cross id, e.g. adid cross docid.</li>\n<li>custom objective function. slight change the lambdarank objective function and directly optimize for the MAP@12. </li>\n<li>not tuning parameters,  just using learning_rate=0.01 num_leaves=127 min_data=1000  num_trees=5000 with early stopping. </li>\n</ol>\n\n<p>I think LightGBM should capable same ability like XGBoost (actually, it always give me better result than XGBoost. And many of our inside projects also prove this). </p>\n\n<p>But maybe LightGBM is not so easy to tune, especially \"num_leaves\". </p>\n\n<p>[/quote]\nThank you very much!!!\nSo with the same feature set, I can easily get 0.68+ with xgboost pairwise rank without tuning but lightgbm only gave about 0.65 ish. I didn't modify the objective function but just wrote a script that saves the model result and evaluates map@12 every 20 rounds, and stops training when map@12 stops improving.</p>\n\n<p>I can think of two reasons that I got much worse scores than you</p>\n\n<p>1)  usually we do feature engineering and only include features that improve the score. I was adding features to xgboost so the final feature set are kind of \"optimized\" for xgboost, but it became much worse with lightgbm</p>\n\n<p>2)  your customized objective function is much better than lambdarank for map@12 or this specific dataset</p>\n\n<p>I think I tried num_leaves = 256 and 128 with other parameters being default, so it shouldn't have been parameter tuning?</p>\n\n<p>It would be great if you can give one or two tips on how to \"better use\" lightgbm :)</p>\n\n<p>[/quote]</p>\n\n<p>Actually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)</p>\n\n<p>I think the most useful parameters are : <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters</a></p>\n\n<p>e.g. limit the min data in one leaves and so on.</p>\n\n<p>And I think another reason is over-fitting. \nLightGBM is more easy to cause over-fitting due to it fits training data better. \nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 157303,
              "author_name": "Little Boat",
              "author_url": "",
              "post_date": "2017-01-20T06:34:38.857000",
              "content": "<p>[quote=insulator;157302]</p>\n\n<p>Actually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)</p>\n\n<p>I think the most useful parameters are : <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters</a></p>\n\n<p>e.g. limit the min data in one leaves and so on.</p>\n\n<p>And I think another reason is over-fitting. \nLightGBM is more easy to cause over-fitting due to it fits training data better. \nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  </p>\n\n<p>[/quote]</p>\n\n<p>That is a good idea! Thanks, will definitely try next time. Will MAP become one of the default evaluation options for lightgbm in future?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 157304,
              "author_name": "Guolin Ke",
              "author_url": "",
              "post_date": "2017-01-20T06:39:15.457000",
              "content": "<p>[quote=Little Boat;157303]</p>\n\n<p>[quote=insulator;157302]</p>\n\n<p>Actually, this custom objective didn't improve much. I can get 0.687+ if just using binary classification objective..( But lambadarank objective give slight bad result than binary ...)</p>\n\n<p>I think the most useful parameters are : <a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters\">https://github.com/Microsoft/LightGBM/blob/master/docs/Parameters.md#learning-control-parameters</a></p>\n\n<p>e.g. limit the min data in one leaves and so on.</p>\n\n<p>And I think another reason is over-fitting. \nLightGBM is more easy to cause over-fitting due to it fits training data better. \nTo solve this, you can output the training loss both for the XGBoost and LightGBM. \nIf the gap between training and validation loss of LightGBM is much larger than XGBoost, you can try to reduce the model complexity of LightGBM to solve this.  </p>\n\n<p>[/quote]</p>\n\n<p>That is a good idea! Thanks, will definitely try next time. Will MAP become one of the default evaluation options for lightgbm in future?</p>\n\n<p>[/quote]</p>\n\n<p>Sure, </p>\n\n<p>BTW, if you met any parameter tuning issues or need new features,  welcome to open issues on github. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 157305,
              "author_name": "Little Boat",
              "author_url": "",
              "post_date": "2017-01-20T06:40:15.023000",
              "content": "<p>[quote=insulator;157304]</p>\n\n<p>Sure, </p>\n\n<p>BTW, if you met any parameter tuning issues or need new features,  welcome to open issues on github. </p>\n\n<p>[/quote]</p>\n\n<p>Definitely will!!</p>\n\n<p>Thanks again! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 154849,
      "author_name": "Sameh Faidi",
      "author_url": "",
      "post_date": "2017-01-08T10:48:50.273000",
      "content": "<p>am I the only one using xgboost? :) untuned single model with about 30 features LB 0.677</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 144947,
      "author_name": "HenryVu",
      "author_url": "",
      "post_date": "2016-11-16T03:21:51.983000",
      "content": "<p>BPR: 0.64830, FM using 10 features + leak: 0.67539</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 144731,
      "author_name": "rokh",
      "author_url": "",
      "post_date": "2016-11-15T07:53:23.217000",
      "content": "<p>0.67928  - my current LB score of single ftrl, ~15 features + some interactions</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 157070,
      "author_name": "Sangxia",
      "author_url": "",
      "post_date": "2017-01-19T03:46:17.737000",
      "content": "<p>@Alexey Noskov congrats! that's amazing. what features did you use for the single model and what was the objective?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 155613,
      "author_name": "peixiang",
      "author_url": "",
      "post_date": "2017-01-12T01:25:46.463000",
      "content": "<p>0.674 using single xgboost</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 144566,
      "author_name": "SRK",
      "author_url": "",
      "post_date": "2016-11-14T17:05:52.110000",
      "content": "<p>Thanks Carl. So there is a lot of potential in FFMs then. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 144375,
      "author_name": "SRK",
      "author_url": "",
      "post_date": "2016-11-14T02:45:53.200000",
      "content": "<p>Gilberto, Great score using just three variables.!</p>\n\n<p>Our best FFM scores 0.669 using multiple variables. </p>\n\n<p>I think people are getting far better scores using FFM than FTRL. This is my first time using FFMs. Any help to improve the FFM scores would be highly appreciated.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 155142,
      "author_name": "kos_",
      "author_url": "",
      "post_date": "2017-01-09T20:52:44.067000",
      "content": "<p>I finally managed to complete a run of FFM with leak. Just to confirm everyone's impression, with almost exactly the same features (I had to remove one from FFM), FFM gets a better LB score compared to FTRL.</p>\n\n<ul>\n<li>FTRL: 0.66200</li>\n<li>FFM: 0.66525</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 155208,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-01-10T04:56:35.360000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 155053,
      "author_name": "Andrii Cherednychenko",
      "author_url": "",
      "post_date": "2017-01-09T08:28:47.950000",
      "content": "<p>I am still on single model</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 154797,
      "author_name": "Benjamin Chu",
      "author_url": "",
      "post_date": "2017-01-08T04:16:33.827000",
      "content": "<p>@Guilherme de Oliveira Santos</p>\n\n<p>One of my submission scored 0.6801 is FTRL on 14 features. The FTRL shared in Kernel is awesome but quite slow, so I use <a href=\"https://github.com/jeongyoonlee/Kaggler\">Kaggler</a> package instead (<a href=\"https://github.com/jeongyoonlee/Kaggler#from-source-code\">install from source code</a>). It runs a epoch within 20 mins even when <code>interaction=True</code>. This might help you iterate faster and find a better hyper parameter set.</p>\n\n<p>Edit: Sorry I post the wrong score, 0.6801 is my CV score, its LB is 0.6787</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 148275,
      "author_name": "Eric Couto",
      "author_url": "",
      "post_date": "2016-12-03T22:50:39.393000",
      "content": "<p>My results:</p>\n\n<pre><code>Model           CV      LB\nFTRL + Leak     0.67321 0.67291\nFFM             0.66171 0.66135\nFFM + Leak      0.67636 -------\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 144354,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2016-11-13T23:44:33.073000",
      "content": "<p>FFM, LB: 0.66725 using 2 original features + &quot;leak&quot;</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 157205,
      "author_name": "diaman",
      "author_url": "",
      "post_date": "2017-01-19T17:11:26.670000",
      "content": "<p>@Sangxia we transformed counts with <code>log2(x + 1)</code> transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as <code>log2(x + 1)</code>. But our best ffm was far from 0.69.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 157210,
          "author_name": "Sangxia",
          "author_url": "",
          "post_date": "2017-01-19T17:34:12.287000",
          "content": "<p>[quote=diaman;157205]</p>\n\n<p>@Sangxia we transformed counts with <code>log2(x + 1)</code> transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as <code>log2(x + 1)</code>. But our best ffm was far from 0.69.</p>\n\n<p>[/quote]</p>\n\n<p>I didn't express myself well. What I am seeing in my CV (let's say no leak, only on future data) is that: \nif we only look at rows for users that have less than say 10 rows in page views, the map was about 0.66; \nas this count increases, the map goes up to above 0.7; but then for the users that have more than 70 rows, it rapidly drops to something like 0.63 or lower. Part of my features is a normalized vector of document ids that a user viewed in the page view file. I was wondering if others observed something like that. Hope this makes some sense.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 157200,
      "author_name": "Sangxia",
      "author_url": "",
      "post_date": "2017-01-19T16:42:30.763000",
      "content": "<p>For teams using FFM, have you looked at your CV map@12 by user page view count? Asking because I noticed mine has a huge drop when the count goes above about 60, and was wondering whether I made a somewhat suboptimal choice of model.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 157198,
      "author_name": "FengLi",
      "author_url": "",
      "post_date": "2017-01-19T16:33:30.430000",
      "content": "<p>@insulator Could you please share the parameters for your LightGBM? Did you include original categorical features or you just transform all of them into CTR? If you could elaborate the features you generate, that'll be great. Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 155640,
      "author_name": "Yehoshaphat Schellekens",
      "author_url": "",
      "post_date": "2017-01-12T06:22:50.830000",
      "content": "<p>@peixiang, after this competition ends, could you please publish your Xgboost source code?</p>\n\n<p>I didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 155862,
          "author_name": "peixiang",
          "author_url": "",
          "post_date": "2017-01-13T02:30:55.273000",
          "content": "<p>[quote=Yehoshaphat Schellekens;155640]</p>\n\n<p>@peixiang, after this competition ends, could you please publish your Xgboost source code?</p>\n\n<p>I didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)</p>\n\n<p>[/quote]\nActually I did not use one hot encoding to encode categorical data. For example, to represent document's topic, there are 300 topics which means normally I need a vector of 300 to represent one doc's topic. Instead, I just use the \"main\" topic of this document as a feature. Determining the \"main\" topic of each document is kind of based on some simple data analysis and heuristics. Maybe a more accurate \"main\" topic or adding an extra \"main\" topic will improve the score.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 155255,
      "author_name": "kos_",
      "author_url": "",
      "post_date": "2017-01-10T09:53:46.087000",
      "content": "<p>@infante, FFM stands for Field-aware Factorization Machines. What I am calling FTRL is SRK's kernel (<a href=\"https://www.kaggle.com/sudalairajkumar/outbrain-click-prediction/ftrl-starter-with-leakage-vars\">https://www.kaggle.com/sudalairajkumar/outbrain-click-prediction/ftrl-starter-with-leakage-vars</a>). </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 155095,
      "author_name": "Xiaofeifei",
      "author_url": "",
      "post_date": "2017-01-09T15:01:00.320000",
      "content": "<p>@Grzegorz Sionkowski can you share how many iteration did you do</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 155094,
      "author_name": "FengLi",
      "author_url": "",
      "post_date": "2017-01-09T14:54:21.610000",
      "content": "<p>@Grzegorz Sionkowski Could you share your method after comp? I'm very interested in how you do that.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 155050,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2017-01-09T08:13:40.967000",
      "content": "<p>0.650 - single model with 4 original features, no leak, no sophisticated ML method, just simple iterations. </p>\n\n<p>EDIT: 17 iterations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154922,
      "author_name": "FengLi",
      "author_url": "",
      "post_date": "2017-01-08T19:44:19.163000",
      "content": "<p>@Sameh Faidi You do some transformation to category variables right? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154874,
      "author_name": "Guilherme de Oliveira Santos",
      "author_url": "",
      "post_date": "2017-01-08T13:43:18.960000",
      "content": "<p>@Benjamin Chu Thanks for the reply! I was wondering about model limitations. Right now I'm at LB 0.67222 using 27 features. My model runs 2 epoch in 40 minutes with peak memory usage of 7GB using the FTRL script by tinrtgu.</p>\n\n<p>I was just about to give up the competition my Notebook has only 12GB but I'll try to reach 0.68x now.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 155207,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-01-10T04:54:08.393000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 154860,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T12:08:25.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154851,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T11:02:48.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154817,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T06:30:54.577000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154813,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T05:34:46.223000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154808,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T05:09:38.903000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154799,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T04:27:12.773000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154698,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-07T18:14:48.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 153036,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-29T15:40:23.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 152466,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-26T14:00:19.923000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 152465,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-26T13:55:31.393000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 152463,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-26T13:47:23.160000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 152461,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-12-26T13:43:27.923000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 155409,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-11T03:18:38.743000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 154942,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-01-08T23:36:21.393000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "144296": "A classic thread.!\r\n\r\nTo start with, our current LB score (0.67455) is from a single FTRL model.",
    "157040": "So, we've just checked our best single and it got 0.70017 Public / 0.70053 Private. \r\n\r\nIt's a 2 times bagged on-disk FFM (my own implementation), which was trained for several hours and used less then 8Gb of memory.",
    "144425": "FFM 0.689 using 24 variables.",
    "157185": "@Sangxia\r\n\r\nObjective was just logloss.\r\n\r\nFeatures - a lot:\r\n\r\n* All what can be joined to clicks from other tables - document, its categories, topics, source, publisher, same for ad doc, platform, location, campaign, advertiser, uid, weekday, hour\r\n* Different counters on prev/future user behaviour - did he viewed this ad, ad with the same doc, ad with doc from the same source, and etc, did he clicked it, same for future\r\n* Information from page views - leak, documents user viewed, document sources user viewed, documents he viewed in one hour after click\r\n* Some special features like diff between ad document creation time and view time, diff between ad document creation time and view document creation time\r\n\r\nMaybe I missed something for now, but we'll write more thorough post with our solution later and share the code.",
    "157102": "I was pretty much on single ffm model, just averaged last 5 computations with slightly different features. Devil is in details though, will write about my approach in solutions sharing thread",
    "156248": "@Benjamin Chu: yes i am using rank pairwise objective and map@12 as eval metric. just remember to set the group when contrusting the dmatrix. You can also get similar performance using simple binary classification. my best single xgb is now scoreing about 0.684 in LB (local cv score 0.688).  i am not doing anything special with categorical features as most of them can be left as they are (numerical presentation). ",
    "155129": "My best single models scores on LB:\r\n\r\nFFM with 21 categorical features + leak                               = 0.67932\r\nLightGBM (lambdarank) with 74 numeric features + leak = 0.6707\r\n\r\nTrying to emsemble LightGBM, and some weaker models (XGBoost and RankLib) with FFM many ways (many types scores and ranks averages, Logistic Regression, GBDT leafs index and score bins), without being able to beat my single FFM model up to now.",
    "157101": "single [LightGBM][1]  model with some counting and CTR features get 0.69019 in PLB. \r\n\r\n\r\n  [1]: https://github.com/Microsoft/LightGBM",
    "154849": "am I the only one using xgboost? :) untuned single model with about 30 features LB 0.677",
    "144947": "BPR: 0.64830, FM using 10 features + leak: 0.67539",
    "144731": "0.67928  - my current LB score of single ftrl, ~15 features + some interactions",
    "157070": "@Alexey Noskov congrats! that's amazing. what features did you use for the single model and what was the objective?",
    "155613": "0.674 using single xgboost",
    "144566": "Thanks Carl. So there is a lot of potential in FFMs then. ",
    "144375": "Gilberto, Great score using just three variables.!\r\n\r\nOur best FFM scores 0.669 using multiple variables. \r\n\r\nI think people are getting far better scores using FFM than FTRL. This is my first time using FFMs. Any help to improve the FFM scores would be highly appreciated.",
    "155142": "I finally managed to complete a run of FFM with leak. Just to confirm everyone's impression, with almost exactly the same features (I had to remove one from FFM), FFM gets a better LB score compared to FTRL.\r\n\r\n- FTRL: 0.66200\r\n- FFM: 0.66525",
    "155053": "I am still on single model",
    "154797": "@Guilherme de Oliveira Santos\r\n\r\nOne of my submission scored 0.6801 is FTRL on 14 features. The FTRL shared in Kernel is awesome but quite slow, so I use [Kaggler][1] package instead ([install from source code][2]). It runs a epoch within 20 mins even when `interaction=True`. This might help you iterate faster and find a better hyper parameter set.\r\n\r\nEdit: Sorry I post the wrong score, 0.6801 is my CV score, its LB is 0.6787\r\n\r\n  [1]: https://github.com/jeongyoonlee/Kaggler\r\n  [2]: https://github.com/jeongyoonlee/Kaggler#from-source-code",
    "148275": "My results:\r\n\r\n    Model           CV  \tLB\r\n    FTRL + Leak \t0.67321\t0.67291\r\n    FFM         \t0.66171\t0.66135\r\n    FFM + Leak      0.67636 -------\r\n\r\n ",
    "144354": "FFM, LB: 0.66725 using 2 original features + \"leak\"",
    "157205": "@Sangxia we transformed counts with `log2(x + 1)` transformation. Actually coding them with 1 (viewed/not viewed) worked pretty the same as `log2(x + 1)`. But our best ffm was far from 0.69.",
    "157200": "For teams using FFM, have you looked at your CV map@12 by user page view count? Asking because I noticed mine has a huge drop when the count goes above about 60, and was wondering whether I made a somewhat suboptimal choice of model.",
    "157198": "@insulator Could you please share the parameters for your LightGBM? Did you include original categorical features or you just transform all of them into CTR? If you could elaborate the features you generate, that'll be great. Thank you.",
    "155640": "@peixiang, after this competition ends, could you please publish your Xgboost source code?\r\n\r\nI didn't manage to efficiently 'one hot encode' the Xgboost input data (there are so many categorical variables)  would love to see how you've done it :)",
    "155255": "@infante, FFM stands for Field-aware Factorization Machines. What I am calling FTRL is SRK's kernel (https://www.kaggle.com/sudalairajkumar/outbrain-click-prediction/ftrl-starter-with-leakage-vars). ",
    "155095": "@Grzegorz Sionkowski can you share how many iteration did you do",
    "155094": "@Grzegorz Sionkowski Could you share your method after comp? I'm very interested in how you do that.",
    "155050": "0.650 - single model with 4 original features, no leak, no sophisticated ML method, just simple iterations. \r\n\r\nEDIT: 17 iterations",
    "154922": "@Sameh Faidi You do some transformation to category variables right? ",
    "154874": "@Benjamin Chu Thanks for the reply! I was wondering about model limitations. Right now I'm at LB 0.67222 using 27 features. My model runs 2 epoch in 40 minutes with peak memory usage of 7GB using the FTRL script by tinrtgu.\r\n\r\nI was just about to give up the competition my Notebook has only 12GB but I'll try to reach 0.68x now.",
    "154860": "@Sameh Faidi\r\n\r\nDo you use pairwise ranking of XGB?",
    "154851": "@Sameh: You're the first I see using xgboost! I was trying to use it last week but with my limited amount of RAM (16Go), I could not run it my many features so I switched to FFM. And even now I still not able to run it on whole train set with more than 10 features...\r\n\r\nYou said you had more than 100Go of RAM, what is your maximum usage? Are you running your model on Python or R?\r\n\r\n",
    "154817": "@-puli-\r\n14 features will have peak memory usage about 21GB. My server's memory is larger than that. I think you may achieve 0.67X with only 5 features, and in that case the data could fit in a 16GB PC.",
    "154813": "@Benjamin Chu: That's a nice script. Just wondering ain't loading the entire dataset in libsvm format heavy on computation. What's your memory like.. just curious :)\r\n",
    "154808": "@FengLi \r\nYep, I think it's the key to winning this competition.",
    "154799": "@Benjamin Chu Cool!. You must do a lot of feature engineering.",
    "154698": "Does anyone managed to score above 0.68 using FTRL?",
    "153036": "Hi all,\r\nBecause of the limitation of the machine, I'm wondering what's the best single model performance without using page_view.csv (except the leak) ? \r\n",
    "152466": "@GuoxinLi Please refer the installation instructions and make sure you have all the requirements fulfilled https://github.com/guestwalk/libffm/blob/master/README",
    "152465": "@Keerath Jaggi Thank you for your reply, I remove \r\n\r\nDFLAG += -DUSEOMP, CXXFLAGS += -fopenmpnow \r\n\r\nnow I get error\r\n\r\n../sdk/graphlab/cppipc/server/comm_server.hpp:22:10: fatal error: \r\n      'nanosockets/socket_errors.hpp' file not found\r\ninclude  nanosockets/socket_errors.hpp\r\n         ^\r\n",
    "152463": "@GuoxinLi\r\nOpenMP is required by FFM for parallelization.\r\nIf OpenMP is not available on your\r\nplatform, then please comment out the following lines in Makefile.\r\n\r\n    DFLAG += -DUSEOMP\r\n    CXXFLAGS += -fopenmp\r\nand then try 'make clean all'",
    "152461": "I have problem to install FFM on my machine, I try to install libffm from https://github.com/turi-code/python-libffm, but it gave me error \" clang: error: unsupported option '-fopenmp' \", does anyone know how to solve it.",
    "155409": "",
    "154942": "@Benjamin Chu: Thanks a lot. "
  }
}