{
  "id": 21246,
  "title": "Best XGBoost Submission you've made?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21246",
  "author_name": "",
  "post_date": "2016-05-26T16:30:44.607Z",
  "votes": null,
  "comment_count": 31,
  "views": 4383,
  "content": "<p>What's the best XGBoost submission you've made? Mine's at <strong>0.25328</strong></p>\n\n<p>0.25328 is of absolutely no use to me at all - and that took hours and hours to compute and used tons of features. 1 tree per category. Slow learning.  Is there more to hope for, or is XGBoost a dead end?</p>",
  "messages": [
    {
      "id": "121468",
      "postDate": "05/26/2016 16:30:44",
      "content": "<p>What's the best XGBoost submission you've made? Mine's at <strong>0.25328</strong></p>\n\n<p>0.25328 is of absolutely no use to me at all - and that took hours and hours to compute and used tons of features. 1 tree per category. Slow learning.  Is there more to hope for, or is XGBoost a dead end?</p>",
      "rawMarkdown": "What's the best XGBoost submission you've made? Mine's at **0.25328**\r\n\r\n\r\n0.25328 is of absolutely no use to me at all - and that took hours and hours to compute and used tons of features. 1 tree per category. Slow learning.  Is there more to hope for, or is XGBoost a dead end?",
      "votes": null
    },
    {
      "id": "121517",
      "postDate": "05/27/2016 01:28:10",
      "content": "<p>Mine is 0.21 using only 23 features in xgboost and i sampled 1% of training data, looks disappointing.</p>",
      "rawMarkdown": "Mine is 0.21 using only 23 features in xgboost and i sampled 1% of training data, looks disappointing.",
      "votes": null
    },
    {
      "id": "121522",
      "postDate": "05/27/2016 03:32:39",
      "content": "<p>@ Mattias Fagerlund</p>\n\n<p>Do you use leaky feature such as rig_destination_distance, user_location_city?  In my previous experiment  (I use random forest, though) it might not be a good idea to train machine learning  model with these features. </p>\n\n<p>Also I noted in my validation result that the performance is better when you train data with is_booking=1. I also try to do weighed sample of click and book data, only to find that the best result is completely drop data with is_booking==0. But I haven't submit my result to lb, so I can't guarantee my observation is correct.</p>",
      "rawMarkdown": "Mattias Fagerlund\r\n\r\nDo you use leaky feature such as rig_destination_distance, user_location_city?  In my previous experiment  (I use random forest, though) it might not be a good idea to train machine learning  model with these features. \r\n\r\nAlso I noted in my validation result that the performance is better when you train data with is_booking=1. I also try to do weighed sample of click and book data, only to find that the best result is completely drop data with is_booking==0. But I haven't submit my result to lb, so I can't guarantee my observation is correct.",
      "votes": null
    },
    {
      "id": "121526",
      "postDate": "05/27/2016 03:37:48",
      "content": "<p>@kuan chen</p>\n\n<p>Yes, I include all features and add a couple of my own. I don't see the model learning very much - or overfitting - so I don't think these features have a negative impact.</p>\n\n<p>I also only use bookings, since there's so many of those to make training difficult enough without looking at clicks.</p>",
      "rawMarkdown": "kuan chen\r\n\r\nYes, I include all features and add a couple of my own. I don't see the model learning very much - or overfitting - so I don't think these features have a negative impact.\r\n\r\nI also only use bookings, since there's so many of those to make training difficult enough without looking at clicks.",
      "votes": null
    },
    {
      "id": "121568",
      "postDate": "05/27/2016 13:55:36",
      "content": "<p>@Mattis\n  I have read about using the data leak, present as almost 33% of the data and then you can train the remaining dataset with your xgboost model. Given your model is returning map 0.25 i guess you could get a lb score of &gt;0.5</p>",
      "rawMarkdown": "Mattis\r\n  I have read about using the data leak, present as almost 33% of the data and then you can train the remaining dataset with your xgboost model. Given your model is returning map 0.25 i guess you could get a lb score of >0.5",
      "votes": null
    },
    {
      "id": "121599",
      "postDate": "05/27/2016 17:58:36",
      "content": "<p>Yep, I'm currently at <strong>#18</strong> with <strong>0.50561</strong>. But that solution didn't use the xgboost solution; at <strong>0.25328</strong> it wasn't good enough to actually help.</p>\n\n<p>/m</p>",
      "rawMarkdown": "Yep, I'm currently at **#18** with **0.50561**. But that solution didn't use the xgboost solution; at **0.25328** it wasn't good enough to actually help.\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "121624",
      "postDate": "05/28/2016 02:31:17",
      "content": "<p>@Mattias Fagerlund  So your current score is based on a rule-based model ? And you run xgboost in your local computer or other ways?  Thank you. </p>",
      "rawMarkdown": "Mattias Fagerlund  So your current score is based on a rule-based model ? And you run xgboost in your local computer or other ways?  Thank you.",
      "votes": null
    },
    {
      "id": "121635",
      "postDate": "05/28/2016 07:12:49",
      "content": "<p>Yes, I evolved a rules-based model similar to the scripts around. Basically I tested millions of models and picked the best one I could generate. That model had the option to include my XGBoost submission, but it wasn't used (probably because it was too weak).</p>",
      "rawMarkdown": "Yes, I evolved a rules-based model similar to the scripts around. Basically I tested millions of models and picked the best one I could generate. That model had the option to include my XGBoost submission, but it wasn't used (probably because it was too weak).",
      "votes": null
    },
    {
      "id": "121660",
      "postDate": "05/28/2016 14:43:29",
      "content": "<p>@Mattias Fagerlund Your XGBoost got 0.25. Did this score include information leak? If not, could you get 0.5+ including data leak using XGBoost model? </p>",
      "rawMarkdown": "Mattias Fagerlund Your XGBoost got 0.25. Did this score include information leak? If not, could you get 0.5+ including data leak using XGBoost model?",
      "votes": null
    },
    {
      "id": "121662",
      "postDate": "05/28/2016 14:51:26",
      "content": "<p>FWIW: best i got from xgboost was 0.30 validation / 0.288 lb, but it took 5h hours on a biiiig machine - and NBayes on a laptop bested it...</p>",
      "rawMarkdown": "FWIW: best i got from xgboost was 0.30 validation / 0.288 lb, but it took 5h hours on a biiiig machine - and NBayes on a laptop bested it...",
      "votes": null
    },
    {
      "id": "121664",
      "postDate": "05/28/2016 14:52:29",
      "content": "<p>@ Mattias Fagerlund</p>\n\n<p>Thank you for your idea, it's my first time to see using genetic algorithm to choose the feature combination, very interesting!</p>",
      "rawMarkdown": "Mattias Fagerlund\r\n\r\nThank you for your idea, it's my first time to see using genetic algorithm to choose the feature combination, very interesting!",
      "votes": null
    },
    {
      "id": "121694",
      "postDate": "05/28/2016 21:42:50",
      "content": "<p>@Fengli  - the data that was the leak was included, but I don't think XGBoost was able to utilize it. But I did get 0.5+ combining my XGBoost model with the data leak. But eventually, it turned out that I did better without the XGBoost model, as outlined here: <a href=\"https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\">https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/</a>.</p>\n\n<p>@Konrad, that's an interesting result. My run took far longer than 5h on a quite beefy machine - possibly not as beefy as yours though. But ran until it stopped learning, so longer time or a bigger machine wouldn't have helped me. Did you use the destinations data in any way? Did you evolve one classifier per hotel cluster? (that's what I did, 100 different xgboost runs). Did you do much feature engineering?</p>\n\n<p>I gotta learn more about NBayes, it would seem! It's now on my &quot;to scratch build&quot; list.</p>\n\n<p>@kuan\nGA can be used to optimize just about anything, though it's rarely the best choice. In this particular case, I found it extremely difficult to build a greedy method that would layer feature selections one on top of the other. The reason why it was hard was that the very best feature combinations to do counting on would be globally terrible (they only gave votes for a small number of samples) but locally great. And my greedy method would have to use some other metric than global effectiveness. </p>\n\n<p>I ended up measuring how well any feature selection did on the data it gave votes on, ignoring everything else. That gave semi-good responses, but it wouldn't find any useful global maxima. </p>\n\n<p>My GA version, on the other hand, measures how a combination of feature selections perform as a whole system - it doesn't attempt to be greedy and add to the solution in increments. That makes it slower, but way easier to code up.</p>",
      "rawMarkdown": "Fengli  - the data that was the leak was included, but I don't think XGBoost was able to utilize it. But I did get 0.5+ combining my XGBoost model with the data leak. But eventually, it turned out that I did better without the XGBoost model, as outlined here: https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/.\r\n\r\n@Konrad, that's an interesting result. My run took far longer than 5h on a quite beefy machine - possibly not as beefy as yours though. But ran until it stopped learning, so longer time or a bigger machine wouldn't have helped me. Did you use the destinations data in any way? Did you evolve one classifier per hotel cluster? (that's what I did, 100 different xgboost runs). Did you do much feature engineering?\r\n\r\nI gotta learn more about NBayes, it would seem! It's now on my \"to scratch build\" list.\r\n\r\n@kuan\r\nGA can be used to optimize just about anything, though it's rarely the best choice. In this particular case, I found it extremely difficult to build a greedy method that would layer feature selections one on top of the other. The reason why it was hard was that the very best feature combinations to do counting on would be globally terrible (they only gave votes for a small number of samples) but locally great. And my greedy method would have to use some other metric than global effectiveness. \r\n\r\nI ended up measuring how well any feature selection did on the data it gave votes on, ignoring everything else. That gave semi-good responses, but it wouldn't find any useful global maxima. \r\n\r\nMy GA version, on the other hand, measures how a combination of feature selections perform as a whole system - it doesn't attempt to be greedy and add to the solution in increments. That makes it slower, but way easier to code up.",
      "votes": null
    },
    {
      "id": "121695",
      "postDate": "05/28/2016 21:46:38",
      "content": "<p>@Mattias: I did use destinations, very minimal feature engineering (purge NA, parse dates etc), multiclass all the way (so technically it's equivalent to a tree per cluster i think, since under the hood xgboost is doing oaa i believe). the machine in question was m4.10xlarge - i decided to roll with it just out of curiosity.</p>",
      "rawMarkdown": "Mattias: I did use destinations, very minimal feature engineering (purge NA, parse dates etc), multiclass all the way (so technically it's equivalent to a tree per cluster i think, since under the hood xgboost is doing oaa i believe). the machine in question was m4.10xlarge - i decided to roll with it just out of curiosity.",
      "votes": null
    },
    {
      "id": "121696",
      "postDate": "05/28/2016 22:00:24",
      "content": "<p>@Konrad: That is a big machine! Wait, did you use all the values from destination as raw features? Or did you pass it through some PCA first?</p>\n\n<p>/m</p>",
      "rawMarkdown": "Konrad: That is a big machine! Wait, did you use all the values from destination as raw features? Or did you pass it through some PCA first?\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "121697",
      "postDate": "05/28/2016 22:01:45",
      "content": "<p>first 60 principal components - that amounted to ~99pct of the variation in the data. </p>",
      "rawMarkdown": "first 60 principal components - that amounted to ~99pct of the variation in the data.",
      "votes": null
    },
    {
      "id": "121698",
      "postDate": "05/28/2016 22:12:38",
      "content": "<p>@Mattias Fagerlund Thank you. Why don't you try small dataset first and find good combination and then apply it to the whole dataset?</p>",
      "rawMarkdown": "Mattias Fagerlund Thank you. Why don't you try small dataset first and find good combination and then apply it to the whole dataset?",
      "votes": null
    },
    {
      "id": "121732",
      "postDate": "05/29/2016 07:28:08",
      "content": "<p>@FengLi, that's what I do, training on the full dataset is hard. With my evolved solutions, I try 10%, and if the performance is acceptable (&gt;80% of the best solution so far), I do a full check.</p>\n\n<p>For XGBoost; I use only random 30% of the data for each iteration.</p>",
      "rawMarkdown": "FengLi, that's what I do, training on the full dataset is hard. With my evolved solutions, I try 10%, and if the performance is acceptable (>80% of the best solution so far), I do a full check.\r\n\r\nFor XGBoost; I use only random 30% of the data for each iteration.",
      "votes": null
    },
    {
      "id": "121750",
      "postDate": "05/29/2016 12:21:54",
      "content": "<p>@Mattias Fagerlund - which part of the train do you use, in the 10%? I mean, how do you choose the 10%?\nAnd in which part are you validating?</p>",
      "rawMarkdown": "Mattias Fagerlund - which part of the train do you use, in the 10%? I mean, how do you choose the 10%?\r\nAnd in which part are you validating?",
      "votes": null
    },
    {
      "id": "121751",
      "postDate": "05/29/2016 12:29:13",
      "content": "<p>Well, I've computed a &quot;submission&quot; on the validation data for each combination of strategies. I check the performance of a combination of combinations of combinations of strategies using 10% of the validation data. If that's good, then I run with the full validation data to compute the full validation result. Whichever gives the best validation result, that's the one I use.</p>\n\n<p>/m</p>",
      "rawMarkdown": "Well, I've computed a \"submission\" on the validation data for each combination of strategies. I check the performance of a combination of combinations of combinations of strategies using 10% of the validation data. If that's good, then I run with the full validation data to compute the full validation result. Whichever gives the best validation result, that's the one I use.\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "121753",
      "postDate": "05/29/2016 12:36:10",
      "content": "<p>How do you choose the first 10%?</p>",
      "rawMarkdown": "How do you choose the first 10%?",
      "votes": null
    },
    {
      "id": "121755",
      "postDate": "05/29/2016 12:38:48",
      "content": "<p>Mmmm. I use c#, so I doubt it will be useful to you, but;</p>\n\n<pre><code>List&lt;Search&gt; shortValidation = validation.Take((int)(ValidationCount * 0.1f)).ToList();\n</code></pre>\n\n<p>But before that, I've randomized the ordering of the validation data so I get a better sample.</p>",
      "rawMarkdown": "Mmmm. I use c#, so I doubt it will be useful to you, but;\r\n       \r\n\r\n    List<Search> shortValidation = validation.Take((int)(ValidationCount * 0.1f)).ToList();\r\n\r\nBut before that, I've randomized the ordering of the validation data so I get a better sample.",
      "votes": null
    },
    {
      "id": "121762",
      "postDate": "05/29/2016 14:45:13",
      "content": "<p>[quote=dot277;121753]</p>\n\n<p>How do you choose the first 10%?</p>\n\n<p>[/quote]</p>\n\n<p>awk -F&quot;,&quot; '$8%10==7 {print}' train.csv</p>",
      "rawMarkdown": "[quote=dot277;121753]\r\n\r\nHow do you choose the first 10%?\r\n\r\n[/quote]\r\n\r\nawk -F\",\" '$8%10==7 {print}' train.csv",
      "votes": null
    },
    {
      "id": "121782",
      "postDate": "05/29/2016 17:11:40",
      "content": "<p>@ Konrad Banachewicz   How long it takes you to run Naive Bayes on your laptop?</p>",
      "rawMarkdown": "Konrad Banachewicz   How long it takes you to run Naive Bayes on your laptop?",
      "votes": null
    },
    {
      "id": "122059",
      "postDate": "06/01/2016 04:48:10",
      "content": "<p>My best xgboost submission is 0.49869. </p>\n\n<p>[Edit] That's actually a blend of ~20 xgboost runs using 1,000,000 training rows each. Best single run so far was 0.48685.</p>",
      "rawMarkdown": "My best xgboost submission is 0.49869. \r\n\r\n[Edit] That's actually a blend of ~20 xgboost runs using 1,000,000 training rows each. Best single run so far was 0.48685.",
      "votes": null
    },
    {
      "id": "122063",
      "postDate": "06/01/2016 05:02:04",
      "content": "<p>20 xgboost using 1000000.  20*1000000&lt;37000000. You didn't use out the full data set? </p>",
      "rawMarkdown": "20 xgboost using 1000000.  20*1000000<37000000. You didn't use out the full data set?",
      "votes": null
    },
    {
      "id": "122066",
      "postDate": "06/01/2016 05:08:49",
      "content": "<p>Nope, not yet at least.</p>",
      "rawMarkdown": "Nope, not yet at least.",
      "votes": null
    },
    {
      "id": "122081",
      "postDate": "06/01/2016 06:33:19",
      "content": "<p>what is your approach to blending?</p>",
      "rawMarkdown": "what is your approach to blending?",
      "votes": null
    },
    {
      "id": "122084",
      "postDate": "06/01/2016 06:50:02",
      "content": "<p>It's just an unweighted average of the class probabilities, then I take the top 5, nothing fancy.</p>",
      "rawMarkdown": "It's just an unweighted average of the class probabilities, then I take the top 5, nothing fancy.",
      "votes": null
    },
    {
      "id": "122097",
      "postDate": "06/01/2016 07:59:00",
      "content": "<p>@branden that's quite impressive! I must be missing some nice features...</p>",
      "rawMarkdown": "branden that's quite impressive! I must be missing some nice features...",
      "votes": null
    },
    {
      "id": "122141",
      "postDate": "06/01/2016 14:47:39",
      "content": "<p>Or maybe forest size. I remember competition where everybody was doing like 2000-3000 trees and one guy made 20000 which had a significant impact on the score.</p>",
      "rawMarkdown": "Or maybe forest size. I remember competition where everybody was doing like 2000-3000 trees and one guy made 20000 which had a significant impact on the score.",
      "votes": null
    },
    {
      "id": "122159",
      "postDate": "06/01/2016 16:54:40",
      "content": "<p>@ Marcin P&#281;kalski Did you mean nrounds?</p>",
      "rawMarkdown": "Marcin Pękalski Did you mean nrounds?",
      "votes": null
    },
    {
      "id": "122364",
      "postDate": "06/03/2016 12:13:32",
      "content": "<p>My best xgboost so far is 0.23891, quite disapointing, but i am still learning :-)</p>",
      "rawMarkdown": "My best xgboost so far is 0.23891, quite disapointing, but i am still learning :-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 121517,
      "author_name": "masterliu",
      "author_url": "",
      "post_date": "05/27/2016 01:28:10",
      "content": "<p>Mine is 0.21 using only 23 features in xgboost and i sampled 1% of training data, looks disappointing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121522,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/27/2016 03:32:39",
      "content": "<p>@ Mattias Fagerlund</p>\n\n<p>Do you use leaky feature such as rig_destination_distance, user_location_city?  In my previous experiment  (I use random forest, though) it might not be a good idea to train machine learning  model with these features. </p>\n\n<p>Also I noted in my validation result that the performance is better when you train data with is_booking=1. I also try to do weighed sample of click and book data, only to find that the best result is completely drop data with is_booking==0. But I haven't submit my result to lb, so I can't guarantee my observation is correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121526,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/27/2016 03:37:48",
      "content": "<p>@kuan chen</p>\n\n<p>Yes, I include all features and add a couple of my own. I don't see the model learning very much - or overfitting - so I don't think these features have a negative impact.</p>\n\n<p>I also only use bookings, since there's so many of those to make training difficult enough without looking at clicks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121568,
      "author_name": "abhinavs01858",
      "author_url": "",
      "post_date": "05/27/2016 13:55:36",
      "content": "<p>@Mattis\n  I have read about using the data leak, present as almost 33% of the data and then you can train the remaining dataset with your xgboost model. Given your model is returning map 0.25 i guess you could get a lb score of &gt;0.5</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121599,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/27/2016 17:58:36",
      "content": "<p>Yep, I'm currently at <strong>#18</strong> with <strong>0.50561</strong>. But that solution didn't use the xgboost solution; at <strong>0.25328</strong> it wasn't good enough to actually help.</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121624,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/28/2016 02:31:17",
      "content": "<p>@Mattias Fagerlund  So your current score is based on a rule-based model ? And you run xgboost in your local computer or other ways?  Thank you. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121635,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/28/2016 07:12:49",
      "content": "<p>Yes, I evolved a rules-based model similar to the scripts around. Basically I tested millions of models and picked the best one I could generate. That model had the option to include my XGBoost submission, but it wasn't used (probably because it was too weak).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121660,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/28/2016 14:43:29",
      "content": "<p>@Mattias Fagerlund Your XGBoost got 0.25. Did this score include information leak? If not, could you get 0.5+ including data leak using XGBoost model? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121662,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/28/2016 14:51:26",
      "content": "<p>FWIW: best i got from xgboost was 0.30 validation / 0.288 lb, but it took 5h hours on a biiiig machine - and NBayes on a laptop bested it...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121664,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/28/2016 14:52:29",
      "content": "<p>@ Mattias Fagerlund</p>\n\n<p>Thank you for your idea, it's my first time to see using genetic algorithm to choose the feature combination, very interesting!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121694,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/28/2016 21:42:50",
      "content": "<p>@Fengli  - the data that was the leak was included, but I don't think XGBoost was able to utilize it. But I did get 0.5+ combining my XGBoost model with the data leak. But eventually, it turned out that I did better without the XGBoost model, as outlined here: <a href=\"https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\">https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/</a>.</p>\n\n<p>@Konrad, that's an interesting result. My run took far longer than 5h on a quite beefy machine - possibly not as beefy as yours though. But ran until it stopped learning, so longer time or a bigger machine wouldn't have helped me. Did you use the destinations data in any way? Did you evolve one classifier per hotel cluster? (that's what I did, 100 different xgboost runs). Did you do much feature engineering?</p>\n\n<p>I gotta learn more about NBayes, it would seem! It's now on my &quot;to scratch build&quot; list.</p>\n\n<p>@kuan\nGA can be used to optimize just about anything, though it's rarely the best choice. In this particular case, I found it extremely difficult to build a greedy method that would layer feature selections one on top of the other. The reason why it was hard was that the very best feature combinations to do counting on would be globally terrible (they only gave votes for a small number of samples) but locally great. And my greedy method would have to use some other metric than global effectiveness. </p>\n\n<p>I ended up measuring how well any feature selection did on the data it gave votes on, ignoring everything else. That gave semi-good responses, but it wouldn't find any useful global maxima. </p>\n\n<p>My GA version, on the other hand, measures how a combination of feature selections perform as a whole system - it doesn't attempt to be greedy and add to the solution in increments. That makes it slower, but way easier to code up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121695,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/28/2016 21:46:38",
      "content": "<p>@Mattias: I did use destinations, very minimal feature engineering (purge NA, parse dates etc), multiclass all the way (so technically it's equivalent to a tree per cluster i think, since under the hood xgboost is doing oaa i believe). the machine in question was m4.10xlarge - i decided to roll with it just out of curiosity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121696,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/28/2016 22:00:24",
      "content": "<p>@Konrad: That is a big machine! Wait, did you use all the values from destination as raw features? Or did you pass it through some PCA first?</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121697,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/28/2016 22:01:45",
      "content": "<p>first 60 principal components - that amounted to ~99pct of the variation in the data. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121698,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/28/2016 22:12:38",
      "content": "<p>@Mattias Fagerlund Thank you. Why don't you try small dataset first and find good combination and then apply it to the whole dataset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121732,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/29/2016 07:28:08",
      "content": "<p>@FengLi, that's what I do, training on the full dataset is hard. With my evolved solutions, I try 10%, and if the performance is acceptable (&gt;80% of the best solution so far), I do a full check.</p>\n\n<p>For XGBoost; I use only random 30% of the data for each iteration.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121750,
      "author_name": "dot277",
      "author_url": "",
      "post_date": "05/29/2016 12:21:54",
      "content": "<p>@Mattias Fagerlund - which part of the train do you use, in the 10%? I mean, how do you choose the 10%?\nAnd in which part are you validating?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121751,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/29/2016 12:29:13",
      "content": "<p>Well, I've computed a &quot;submission&quot; on the validation data for each combination of strategies. I check the performance of a combination of combinations of combinations of strategies using 10% of the validation data. If that's good, then I run with the full validation data to compute the full validation result. Whichever gives the best validation result, that's the one I use.</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121753,
      "author_name": "dot277",
      "author_url": "",
      "post_date": "05/29/2016 12:36:10",
      "content": "<p>How do you choose the first 10%?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121755,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/29/2016 12:38:48",
      "content": "<p>Mmmm. I use c#, so I doubt it will be useful to you, but;</p>\n\n<pre><code>List&lt;Search&gt; shortValidation = validation.Take((int)(ValidationCount * 0.1f)).ToList();\n</code></pre>\n\n<p>But before that, I've randomized the ordering of the validation data so I get a better sample.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121762,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "05/29/2016 14:45:13",
      "content": "<p>[quote=dot277;121753]</p>\n\n<p>How do you choose the first 10%?</p>\n\n<p>[/quote]</p>\n\n<p>awk -F&quot;,&quot; '$8%10==7 {print}' train.csv</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121782,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/29/2016 17:11:40",
      "content": "<p>@ Konrad Banachewicz   How long it takes you to run Naive Bayes on your laptop?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122059,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "06/01/2016 04:48:10",
      "content": "<p>My best xgboost submission is 0.49869. </p>\n\n<p>[Edit] That's actually a blend of ~20 xgboost runs using 1,000,000 training rows each. Best single run so far was 0.48685.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122063,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/01/2016 05:02:04",
      "content": "<p>20 xgboost using 1000000.  20*1000000&lt;37000000. You didn't use out the full data set? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122066,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "06/01/2016 05:08:49",
      "content": "<p>Nope, not yet at least.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122081,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "06/01/2016 06:33:19",
      "content": "<p>what is your approach to blending?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122084,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "06/01/2016 06:50:02",
      "content": "<p>It's just an unweighted average of the class probabilities, then I take the top 5, nothing fancy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122097,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/01/2016 07:59:00",
      "content": "<p>@branden that's quite impressive! I must be missing some nice features...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122141,
      "author_name": "mpekalski",
      "author_url": "",
      "post_date": "06/01/2016 14:47:39",
      "content": "<p>Or maybe forest size. I remember competition where everybody was doing like 2000-3000 trees and one guy made 20000 which had a significant impact on the score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122159,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/01/2016 16:54:40",
      "content": "<p>@ Marcin P&#281;kalski Did you mean nrounds?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122364,
      "author_name": "savioz",
      "author_url": "",
      "post_date": "06/03/2016 12:13:32",
      "content": "<p>My best xgboost so far is 0.23891, quite disapointing, but i am still learning :-)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "121468": "What's the best XGBoost submission you've made? Mine's at **0.25328**\r\n\r\n\r\n0.25328 is of absolutely no use to me at all - and that took hours and hours to compute and used tons of features. 1 tree per category. Slow learning.  Is there more to hope for, or is XGBoost a dead end?",
    "121517": "Mine is 0.21 using only 23 features in xgboost and i sampled 1% of training data, looks disappointing.",
    "121522": "Mattias Fagerlund\r\n\r\nDo you use leaky feature such as rig_destination_distance, user_location_city?  In my previous experiment  (I use random forest, though) it might not be a good idea to train machine learning  model with these features. \r\n\r\nAlso I noted in my validation result that the performance is better when you train data with is_booking=1. I also try to do weighed sample of click and book data, only to find that the best result is completely drop data with is_booking==0. But I haven't submit my result to lb, so I can't guarantee my observation is correct.",
    "121526": "kuan chen\r\n\r\nYes, I include all features and add a couple of my own. I don't see the model learning very much - or overfitting - so I don't think these features have a negative impact.\r\n\r\nI also only use bookings, since there's so many of those to make training difficult enough without looking at clicks.",
    "121568": "Mattis\r\n  I have read about using the data leak, present as almost 33% of the data and then you can train the remaining dataset with your xgboost model. Given your model is returning map 0.25 i guess you could get a lb score of >0.5",
    "121599": "Yep, I'm currently at **#18** with **0.50561**. But that solution didn't use the xgboost solution; at **0.25328** it wasn't good enough to actually help.\r\n\r\n/m",
    "121624": "Mattias Fagerlund  So your current score is based on a rule-based model ? And you run xgboost in your local computer or other ways?  Thank you.",
    "121635": "Yes, I evolved a rules-based model similar to the scripts around. Basically I tested millions of models and picked the best one I could generate. That model had the option to include my XGBoost submission, but it wasn't used (probably because it was too weak).",
    "121660": "Mattias Fagerlund Your XGBoost got 0.25. Did this score include information leak? If not, could you get 0.5+ including data leak using XGBoost model?",
    "121662": "FWIW: best i got from xgboost was 0.30 validation / 0.288 lb, but it took 5h hours on a biiiig machine - and NBayes on a laptop bested it...",
    "121664": "Mattias Fagerlund\r\n\r\nThank you for your idea, it's my first time to see using genetic algorithm to choose the feature combination, very interesting!",
    "121694": "Fengli  - the data that was the leak was included, but I don't think XGBoost was able to utilize it. But I did get 0.5+ combining my XGBoost model with the data leak. But eventually, it turned out that I did better without the XGBoost model, as outlined here: https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/.\r\n\r\n@Konrad, that's an interesting result. My run took far longer than 5h on a quite beefy machine - possibly not as beefy as yours though. But ran until it stopped learning, so longer time or a bigger machine wouldn't have helped me. Did you use the destinations data in any way? Did you evolve one classifier per hotel cluster? (that's what I did, 100 different xgboost runs). Did you do much feature engineering?\r\n\r\nI gotta learn more about NBayes, it would seem! It's now on my \"to scratch build\" list.\r\n\r\n@kuan\r\nGA can be used to optimize just about anything, though it's rarely the best choice. In this particular case, I found it extremely difficult to build a greedy method that would layer feature selections one on top of the other. The reason why it was hard was that the very best feature combinations to do counting on would be globally terrible (they only gave votes for a small number of samples) but locally great. And my greedy method would have to use some other metric than global effectiveness. \r\n\r\nI ended up measuring how well any feature selection did on the data it gave votes on, ignoring everything else. That gave semi-good responses, but it wouldn't find any useful global maxima. \r\n\r\nMy GA version, on the other hand, measures how a combination of feature selections perform as a whole system - it doesn't attempt to be greedy and add to the solution in increments. That makes it slower, but way easier to code up.",
    "121695": "Mattias: I did use destinations, very minimal feature engineering (purge NA, parse dates etc), multiclass all the way (so technically it's equivalent to a tree per cluster i think, since under the hood xgboost is doing oaa i believe). the machine in question was m4.10xlarge - i decided to roll with it just out of curiosity.",
    "121696": "Konrad: That is a big machine! Wait, did you use all the values from destination as raw features? Or did you pass it through some PCA first?\r\n\r\n/m",
    "121697": "first 60 principal components - that amounted to ~99pct of the variation in the data.",
    "121698": "Mattias Fagerlund Thank you. Why don't you try small dataset first and find good combination and then apply it to the whole dataset?",
    "121732": "FengLi, that's what I do, training on the full dataset is hard. With my evolved solutions, I try 10%, and if the performance is acceptable (>80% of the best solution so far), I do a full check.\r\n\r\nFor XGBoost; I use only random 30% of the data for each iteration.",
    "121750": "Mattias Fagerlund - which part of the train do you use, in the 10%? I mean, how do you choose the 10%?\r\nAnd in which part are you validating?",
    "121751": "Well, I've computed a \"submission\" on the validation data for each combination of strategies. I check the performance of a combination of combinations of combinations of strategies using 10% of the validation data. If that's good, then I run with the full validation data to compute the full validation result. Whichever gives the best validation result, that's the one I use.\r\n\r\n/m",
    "121753": "How do you choose the first 10%?",
    "121755": "Mmmm. I use c#, so I doubt it will be useful to you, but;\r\n       \r\n\r\n    List<Search> shortValidation = validation.Take((int)(ValidationCount * 0.1f)).ToList();\r\n\r\nBut before that, I've randomized the ordering of the validation data so I get a better sample.",
    "121762": "[quote=dot277;121753]\r\n\r\nHow do you choose the first 10%?\r\n\r\n[/quote]\r\n\r\nawk -F\",\" '$8%10==7 {print}' train.csv",
    "121782": "Konrad Banachewicz   How long it takes you to run Naive Bayes on your laptop?",
    "122059": "My best xgboost submission is 0.49869. \r\n\r\n[Edit] That's actually a blend of ~20 xgboost runs using 1,000,000 training rows each. Best single run so far was 0.48685.",
    "122063": "20 xgboost using 1000000.  20*1000000<37000000. You didn't use out the full data set?",
    "122066": "Nope, not yet at least.",
    "122081": "what is your approach to blending?",
    "122084": "It's just an unweighted average of the class probabilities, then I take the top 5, nothing fancy.",
    "122097": "branden that's quite impressive! I must be missing some nice features...",
    "122141": "Or maybe forest size. I remember competition where everybody was doing like 2000-3000 trees and one guy made 20000 which had a significant impact on the score.",
    "122159": "Marcin Pękalski Did you mean nrounds?",
    "122364": "My best xgboost so far is 0.23891, quite disapointing, but i am still learning :-)"
  },
  "source": "meta"
}