{
  "id": 20831,
  "title": "Why random forest does not work well?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20831",
  "author_name": "",
  "post_date": "2016-05-10T08:29:48.940Z",
  "votes": 2,
  "comment_count": 21,
  "views": 4024,
  "content": "<p>I currently work my model based on sklearn random forest. I select several features: orig_destination_distance, usr_location_city, srch_destination_id,hotel_market. \nI randomly sample, 10000 data from training set and train a random forest then aggregate the prediction of each forest.</p>\n\n<p>Though there are lots of categorical data in the training set, I read through <a href=\"https://www.mail-archive.com/scikit-learn-general@lists.sourceforge.net/msg07373.html\">this mailing list</a> say that it still possible to apply random forest to such high cardinality data</p>\n\n<p>What I observe is that the oob score is around 0.35~0.39. However, the performance on test data drop to around 0.22.</p>\n\n<p>I have seen <a href=\"https://www.dataquest.io/blog/kaggle-tutorial/\">this post</a> that say the ML technique might not work but still don't know exactly why? Especially the oob_score look not bad. Is ml technique totally  not suitable for this problem or just because my model might overfit? Thanks a lot</p>",
  "messages": [
    {
      "id": "119434",
      "postDate": "05/10/2016 08:29:48",
      "content": "<p>I currently work my model based on sklearn random forest. I select several features: orig_destination_distance, usr_location_city, srch_destination_id,hotel_market. \nI randomly sample, 10000 data from training set and train a random forest then aggregate the prediction of each forest.</p>\n\n<p>Though there are lots of categorical data in the training set, I read through <a href=\"https://www.mail-archive.com/scikit-learn-general@lists.sourceforge.net/msg07373.html\">this mailing list</a> say that it still possible to apply random forest to such high cardinality data</p>\n\n<p>What I observe is that the oob score is around 0.35~0.39. However, the performance on test data drop to around 0.22.</p>\n\n<p>I have seen <a href=\"https://www.dataquest.io/blog/kaggle-tutorial/\">this post</a> that say the ML technique might not work but still don't know exactly why? Especially the oob_score look not bad. Is ml technique totally  not suitable for this problem or just because my model might overfit? Thanks a lot</p>",
      "rawMarkdown": "I currently work my model based on sklearn random forest. I select several features: orig_destination_distance, usr_location_city, srch_destination_id,hotel_market. \r\nI randomly sample, 10000 data from training set and train a random forest then aggregate the prediction of each forest.\r\n\r\nThough there are lots of categorical data in the training set, I read through [this mailing list][1] say that it still possible to apply random forest to such high cardinality data\r\n\r\nWhat I observe is that the oob score is around 0.35~0.39. However, the performance on test data drop to around 0.22.\r\n\r\nI have seen [this post][2] that say the ML technique might not work but still don't know exactly why? Especially the oob_score look not bad. Is ml technique totally  not suitable for this problem or just because my model might overfit? Thanks a lot\r\n\r\n\r\n  [1]: https://www.mail-archive.com/scikit-learn-general@lists.sourceforge.net/msg07373.html\r\n  [2]: https://www.dataquest.io/blog/kaggle-tutorial/",
      "votes": null
    },
    {
      "id": "119437",
      "postDate": "05/10/2016 08:53:42",
      "content": "<p>Did you try just using 2014 stats?</p>",
      "rawMarkdown": "Did you try just using 2014 stats?",
      "votes": null
    },
    {
      "id": "119440",
      "postDate": "05/10/2016 09:12:36",
      "content": "<p>@Scirpus</p>\n\n<p>No I used whole data set, I saw several script in top 60 that use only 2014 data set. It's interesting I will try it</p>",
      "rawMarkdown": "Scirpus\r\n\r\nNo I used whole data set, I saw several script in top 60 that use only 2014 data set. It's interesting I will try it",
      "votes": null
    },
    {
      "id": "119619",
      "postDate": "05/11/2016 22:20:33",
      "content": "<p>I am little bit worry about your features. orig_destination_distance, usr_location_city, srch_destination_id,hotel_market have been confirmed to have data leak.\nIf you train with those features, I don't think the model is well generalized.</p>",
      "rawMarkdown": "I am little bit worry about your features. orig_destination_distance, usr_location_city, srch_destination_id,hotel_market have been confirmed to have data leak.\r\nIf you train with those features, I don't think the model is well generalized.",
      "votes": null
    },
    {
      "id": "119629",
      "postDate": "05/12/2016 00:41:55",
      "content": "<p>@Li Li. Same problem for XGB. It wastes me so much time until I realize XGB perform better than local most popular guess is because of data leakage, not the model itself.</p>",
      "rawMarkdown": "Li Li. Same problem for XGB. It wastes me so much time until I realize XGB perform better than local most popular guess is because of data leakage, not the model itself.",
      "votes": null
    },
    {
      "id": "119634",
      "postDate": "05/12/2016 01:03:59",
      "content": "<p>@YijunMa, exactly. We wasted a lot of electricities...</p>",
      "rawMarkdown": "YijunMa, exactly. We wasted a lot of electricities...",
      "votes": null
    },
    {
      "id": "119637",
      "postDate": "05/12/2016 01:12:21",
      "content": "<p>@Li Li\nThank you for your suggestion. But I don't quite understand why those data leak feature make the model fail to generalize. Could you explain a little bit?</p>\n\n<p>@YijunMa\nI also try xgboost, I get the similar result as you lol</p>",
      "rawMarkdown": "Li Li\r\nThank you for your suggestion. But I don't quite understand why those data leak feature make the model fail to generalize. Could you explain a little bit?\r\n\r\n@YijunMa\r\nI also try xgboost, I get the similar result as you lol",
      "votes": null
    },
    {
      "id": "119638",
      "postDate": "05/12/2016 01:13:31",
      "content": "<p>@Li Li. I am trying to round up the result, hopefully the leakage problem can be solved. The trick is that distance itself might mean something. Say a long distance might means a international travel, and in this case the cluster could narrow to certain range.</p>\n\n<p>Doing feature engineering in such a big data set is so painful... Do you have any clue now?</p>",
      "rawMarkdown": "Li Li. I am trying to round up the result, hopefully the leakage problem can be solved. The trick is that distance itself might mean something. Say a long distance might means a international travel, and in this case the cluster could narrow to certain range.\r\n\r\nDoing feature engineering in such a big data set is so painful... Do you have any clue now?",
      "votes": null
    },
    {
      "id": "119639",
      "postDate": "05/12/2016 01:31:41",
      "content": "<p>@kuan chen.</p>\n\n<p>What I did is:</p>\n\n<ol>\n<li>5-fold cv, and all the test result is better than local popular around 0.05 at least.</li>\n<li>So excited, and submit my answer (of course updated some of the rows with leakage answers), waiting for top....</li>\n<li>Dala, the result is even worse!!</li>\n<li>I have no idea why, cuz basically it means cross validation is not working. Is it because of my wrong code? Then I realize, aha, I have another hold out test group maybe I can used for testing, which is the leakage data group. I though since it is part of the LB data, it should be a good test set.</li>\n<li>For this test set, the XGB is still better than local popular guess. I am now crazy, and have no idea what to trust...</li>\n<li>Until, I check the feature importance (maybe I should do this earlier). the distance feature is almost the only important feature, and the second important one is user's area.</li>\n<li>As I wipe these two features, XGB performs worse than local popular. Now that I am almost 80% sure it's a data leakage reason.</li>\n</ol>\n\n<p>The idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the &quot;local top 5 popular hotel model&quot; perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.</p>",
      "rawMarkdown": "kuan chen.\r\n\r\nWhat I did is:\r\n\r\n 1. 5-fold cv, and all the test result is better than local popular around 0.05 at least.\r\n 2. So excited, and submit my answer (of course updated some of the rows with leakage answers), waiting for top....\r\n 3. Dala, the result is even worse!!\r\n 4. I have no idea why, cuz basically it means cross validation is not working. Is it because of my wrong code? Then I realize, aha, I have another hold out test group maybe I can used for testing, which is the leakage data group. I though since it is part of the LB data, it should be a good test set.\r\n 5. For this test set, the XGB is still better than local popular guess. I am now crazy, and have no idea what to trust...\r\n 6. Until, I check the feature importance (maybe I should do this earlier). the distance feature is almost the only important feature, and the second important one is user's area.\r\n 7. As I wipe these two features, XGB performs worse than local popular. Now that I am almost 80% sure it's a data leakage reason.\r\n\r\nThe idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the \"local top 5 popular hotel model\" perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.",
      "votes": null
    },
    {
      "id": "119644",
      "postDate": "05/12/2016 01:56:57",
      "content": "<p>@YijunMa</p>\n\n<p>Thank for share your experiment and result! I did the feature importance first by random forest and get the same observation as yours (but at that time, I didn't know about data leakage would cause severe problem on model..., now I learned). </p>\n\n<p>I used data leakage solution now but I am still considering whether it is possible to integrate those models and leakage solution</p>",
      "rawMarkdown": "YijunMa\r\n\r\nThank for share your experiment and result! I did the feature importance first by random forest and get the same observation as yours (but at that time, I didn't know about data leakage would cause severe problem on model..., now I learned). \r\n\r\nI used data leakage solution now but I am still considering whether it is possible to integrate those models and leakage solution",
      "votes": null
    },
    {
      "id": "121495",
      "postDate": "05/26/2016 21:55:47",
      "content": "<p>Hi guys,\nI blend my rfr results with leakage and local popular results. I can't pass lb=0.5\nwhat's your best lb score getting by using xgb or rfr?</p>",
      "rawMarkdown": "Hi guys,\r\nI blend my rfr results with leakage and local popular results. I can't pass lb=0.5\r\nwhat's your best lb score getting by using xgb or rfr?",
      "votes": null
    },
    {
      "id": "121523",
      "postDate": "05/27/2016 03:34:30",
      "content": "<p>I am also curious about the best lb score of xgboost (for me the training the process is too slow...)</p>",
      "rawMarkdown": "I am also curious about the best lb score of xgboost (for me the training the process is too slow...)",
      "votes": null
    },
    {
      "id": "121528",
      "postDate": "05/27/2016 03:40:46",
      "content": "<p>With pure XGBoost, my best score is <strong>0.25328</strong>. But that's to weak to use in my submissions at this point. I've tried to mix it in with my counting strategies, but it gets dropped because it's too weak; <a href=\"https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\">https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/</a> .</p>",
      "rawMarkdown": "With pure XGBoost, my best score is **0.25328**. But that's to weak to use in my submissions at this point. I've tried to mix it in with my counting strategies, but it gets dropped because it's too weak; https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/ .",
      "votes": null
    },
    {
      "id": "121644",
      "postDate": "05/28/2016 09:57:34",
      "content": "<p>@YijunMa - why did you wiped the this two features?\nnot all of the rows are leaky rows, as you know.</p>",
      "rawMarkdown": "YijunMa - why did you wiped the this two features?\r\nnot all of the rows are leaky rows, as you know.",
      "votes": null
    },
    {
      "id": "121646",
      "postDate": "05/28/2016 10:03:11",
      "content": "<p>I didn't exactly wipe it, it just wasn't used by my evolutionary solution; <a href=\"https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\">https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/</a></p>\n\n<p>The thing is, these counters don't only work for leaky data. There are public scripts that do &gt;0.5 by using counters. Ie: pick the hotel clusters that best match some set of features in your search. </p>\n\n<p>There isn't much Machine Learning involved in hand picking these counters, but as my blog post descibes, I was able to create a system that evolves them.</p>\n\n<p>That system had the option to use the XGBoost submission, but it ultimately wasn't used for the best solutions.</p>",
      "rawMarkdown": "I didn't exactly wipe it, it just wasn't used by my evolutionary solution; https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\r\n\r\nThe thing is, these counters don't only work for leaky data. There are public scripts that do >0.5 by using counters. Ie: pick the hotel clusters that best match some set of features in your search. \r\n\r\nThere isn't much Machine Learning involved in hand picking these counters, but as my blog post descibes, I was able to create a system that evolves them.\r\n\r\nThat system had the option to use the XGBoost submission, but it ultimately wasn't used for the best solutions.",
      "votes": null
    },
    {
      "id": "121655",
      "postDate": "05/28/2016 14:16:43",
      "content": "<p>@Mattias Fagerlund Thank you for sharing. Did you use any information of the destination file? </p>",
      "rawMarkdown": "Mattias Fagerlund Thank you for sharing. Did you use any information of the destination file?",
      "votes": null
    },
    {
      "id": "121714",
      "postDate": "05/29/2016 01:44:12",
      "content": "<p>BTW I just ran the .50161 script excluding 2015 training data, and it got .49266.</p>",
      "rawMarkdown": "BTW I just ran the .50161 script excluding 2015 training data, and it got .49266.",
      "votes": null
    },
    {
      "id": "121726",
      "postDate": "05/29/2016 04:43:32",
      "content": "<p>@happycube  what does that mean? Training data just includes 2013 and 2014 as I remembered.</p>",
      "rawMarkdown": "happycube  what does that mean? Training data just includes 2013 and 2014 as I remembered.",
      "votes": null
    },
    {
      "id": "121728",
      "postDate": "05/29/2016 06:02:23",
      "content": "<p>Oops - I meant to say year for checkin date (book_year in the script) - of which about 12-13% is dated 2015.</p>\n\n<p>2013 clicks only gets .40609, 2014 .48472</p>",
      "rawMarkdown": "Oops - I meant to say year for checkin date (book_year in the script) - of which about 12-13% is dated 2015.\r\n\r\n2013 clicks only gets .40609, 2014 .48472",
      "votes": null
    },
    {
      "id": "121910",
      "postDate": "05/30/2016 21:16:46",
      "content": "<p>&quot;The idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the &quot;local top 5 popular hotel model&quot; perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.&quot;</p>\n\n<p>@YijunMa  You mean that &quot;local top 5 popular hotel model&quot; with leakage solution is better than XGB. Here, is this XGB with leakage solution or without?  Thank you.</p>",
      "rawMarkdown": "\"The idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the \"local top 5 popular hotel model\" perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.\"\r\n\r\n@YijunMa  You mean that \"local top 5 popular hotel model\" with leakage solution is better than XGB. Here, is this XGB with leakage solution or without?  Thank you.",
      "votes": null
    },
    {
      "id": "121917",
      "postDate": "05/30/2016 22:48:25",
      "content": "<p>In my tests, I was never able to make XGB perform better than most popular hotels with leakage. I work on a sample of time ordered 10k bookings, and I was barely able to beat pop hotels without leakage, and LB score confirms. It wastes a lot of time for what it performs. As @Mattias, I got better submissions by mixing xgb predictions with different leakage solutions..</p>",
      "rawMarkdown": "In my tests, I was never able to make XGB perform better than most popular hotels with leakage. I work on a sample of time ordered 10k bookings, and I was barely able to beat pop hotels without leakage, and LB score confirms. It wastes a lot of time for what it performs. As @Mattias, I got better submissions by mixing xgb predictions with different leakage solutions..",
      "votes": null
    },
    {
      "id": "121924",
      "postDate": "05/31/2016 01:23:41",
      "content": "<p>@FengLi, my example here is trying to say how leakage data can ruin the XGB model - even worse than local pop5 model. In this case, the XGB model is merged with leakage solution, so that the two models are compare in the same criteria.</p>",
      "rawMarkdown": "FengLi, my example here is trying to say how leakage data can ruin the XGB model - even worse than local pop5 model. In this case, the XGB model is merged with leakage solution, so that the two models are compare in the same criteria.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119437,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/10/2016 08:53:42",
      "content": "<p>Did you try just using 2014 stats?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119440,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/10/2016 09:12:36",
      "content": "<p>@Scirpus</p>\n\n<p>No I used whole data set, I saw several script in top 60 that use only 2014 data set. It's interesting I will try it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119619,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/11/2016 22:20:33",
      "content": "<p>I am little bit worry about your features. orig_destination_distance, usr_location_city, srch_destination_id,hotel_market have been confirmed to have data leak.\nIf you train with those features, I don't think the model is well generalized.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119629,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/12/2016 00:41:55",
      "content": "<p>@Li Li. Same problem for XGB. It wastes me so much time until I realize XGB perform better than local most popular guess is because of data leakage, not the model itself.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119634,
      "author_name": "aikinogard",
      "author_url": "",
      "post_date": "05/12/2016 01:03:59",
      "content": "<p>@YijunMa, exactly. We wasted a lot of electricities...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119637,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/12/2016 01:12:21",
      "content": "<p>@Li Li\nThank you for your suggestion. But I don't quite understand why those data leak feature make the model fail to generalize. Could you explain a little bit?</p>\n\n<p>@YijunMa\nI also try xgboost, I get the similar result as you lol</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119638,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/12/2016 01:13:31",
      "content": "<p>@Li Li. I am trying to round up the result, hopefully the leakage problem can be solved. The trick is that distance itself might mean something. Say a long distance might means a international travel, and in this case the cluster could narrow to certain range.</p>\n\n<p>Doing feature engineering in such a big data set is so painful... Do you have any clue now?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119639,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/12/2016 01:31:41",
      "content": "<p>@kuan chen.</p>\n\n<p>What I did is:</p>\n\n<ol>\n<li>5-fold cv, and all the test result is better than local popular around 0.05 at least.</li>\n<li>So excited, and submit my answer (of course updated some of the rows with leakage answers), waiting for top....</li>\n<li>Dala, the result is even worse!!</li>\n<li>I have no idea why, cuz basically it means cross validation is not working. Is it because of my wrong code? Then I realize, aha, I have another hold out test group maybe I can used for testing, which is the leakage data group. I though since it is part of the LB data, it should be a good test set.</li>\n<li>For this test set, the XGB is still better than local popular guess. I am now crazy, and have no idea what to trust...</li>\n<li>Until, I check the feature importance (maybe I should do this earlier). the distance feature is almost the only important feature, and the second important one is user's area.</li>\n<li>As I wipe these two features, XGB performs worse than local popular. Now that I am almost 80% sure it's a data leakage reason.</li>\n</ol>\n\n<p>The idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the &quot;local top 5 popular hotel model&quot; perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119644,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/12/2016 01:56:57",
      "content": "<p>@YijunMa</p>\n\n<p>Thank for share your experiment and result! I did the feature importance first by random forest and get the same observation as yours (but at that time, I didn't know about data leakage would cause severe problem on model..., now I learned). </p>\n\n<p>I used data leakage solution now but I am still considering whether it is possible to integrate those models and leakage solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121495,
      "author_name": "hswater1",
      "author_url": "",
      "post_date": "05/26/2016 21:55:47",
      "content": "<p>Hi guys,\nI blend my rfr results with leakage and local popular results. I can't pass lb=0.5\nwhat's your best lb score getting by using xgb or rfr?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121523,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/27/2016 03:34:30",
      "content": "<p>I am also curious about the best lb score of xgboost (for me the training the process is too slow...)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121528,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/27/2016 03:40:46",
      "content": "<p>With pure XGBoost, my best score is <strong>0.25328</strong>. But that's to weak to use in my submissions at this point. I've tried to mix it in with my counting strategies, but it gets dropped because it's too weak; <a href=\"https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\">https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/</a> .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121644,
      "author_name": "dot277",
      "author_url": "",
      "post_date": "05/28/2016 09:57:34",
      "content": "<p>@YijunMa - why did you wiped the this two features?\nnot all of the rows are leaky rows, as you know.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121646,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "05/28/2016 10:03:11",
      "content": "<p>I didn't exactly wipe it, it just wasn't used by my evolutionary solution; <a href=\"https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\">https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/</a></p>\n\n<p>The thing is, these counters don't only work for leaky data. There are public scripts that do &gt;0.5 by using counters. Ie: pick the hotel clusters that best match some set of features in your search. </p>\n\n<p>There isn't much Machine Learning involved in hand picking these counters, but as my blog post descibes, I was able to create a system that evolves them.</p>\n\n<p>That system had the option to use the XGBoost submission, but it ultimately wasn't used for the best solutions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121655,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/28/2016 14:16:43",
      "content": "<p>@Mattias Fagerlund Thank you for sharing. Did you use any information of the destination file? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121714,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "05/29/2016 01:44:12",
      "content": "<p>BTW I just ran the .50161 script excluding 2015 training data, and it got .49266.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121726,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/29/2016 04:43:32",
      "content": "<p>@happycube  what does that mean? Training data just includes 2013 and 2014 as I remembered.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121728,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "05/29/2016 06:02:23",
      "content": "<p>Oops - I meant to say year for checkin date (book_year in the script) - of which about 12-13% is dated 2015.</p>\n\n<p>2013 clicks only gets .40609, 2014 .48472</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121910,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "05/30/2016 21:16:46",
      "content": "<p>&quot;The idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the &quot;local top 5 popular hotel model&quot; perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.&quot;</p>\n\n<p>@YijunMa  You mean that &quot;local top 5 popular hotel model&quot; with leakage solution is better than XGB. Here, is this XGB with leakage solution or without?  Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121917,
      "author_name": "aminkh",
      "author_url": "",
      "post_date": "05/30/2016 22:48:25",
      "content": "<p>In my tests, I was never able to make XGB perform better than most popular hotels with leakage. I work on a sample of time ordered 10k bookings, and I was barely able to beat pop hotels without leakage, and LB score confirms. It wastes a lot of time for what it performs. As @Mattias, I got better submissions by mixing xgb predictions with different leakage solutions..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121924,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/31/2016 01:23:41",
      "content": "<p>@FengLi, my example here is trying to say how leakage data can ruin the XGB model - even worse than local pop5 model. In this case, the XGB model is merged with leakage solution, so that the two models are compare in the same criteria.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119434": "I currently work my model based on sklearn random forest. I select several features: orig_destination_distance, usr_location_city, srch_destination_id,hotel_market. \r\nI randomly sample, 10000 data from training set and train a random forest then aggregate the prediction of each forest.\r\n\r\nThough there are lots of categorical data in the training set, I read through [this mailing list][1] say that it still possible to apply random forest to such high cardinality data\r\n\r\nWhat I observe is that the oob score is around 0.35~0.39. However, the performance on test data drop to around 0.22.\r\n\r\nI have seen [this post][2] that say the ML technique might not work but still don't know exactly why? Especially the oob_score look not bad. Is ml technique totally  not suitable for this problem or just because my model might overfit? Thanks a lot\r\n\r\n\r\n  [1]: https://www.mail-archive.com/scikit-learn-general@lists.sourceforge.net/msg07373.html\r\n  [2]: https://www.dataquest.io/blog/kaggle-tutorial/",
    "119437": "Did you try just using 2014 stats?",
    "119440": "Scirpus\r\n\r\nNo I used whole data set, I saw several script in top 60 that use only 2014 data set. It's interesting I will try it",
    "119619": "I am little bit worry about your features. orig_destination_distance, usr_location_city, srch_destination_id,hotel_market have been confirmed to have data leak.\r\nIf you train with those features, I don't think the model is well generalized.",
    "119629": "Li Li. Same problem for XGB. It wastes me so much time until I realize XGB perform better than local most popular guess is because of data leakage, not the model itself.",
    "119634": "YijunMa, exactly. We wasted a lot of electricities...",
    "119637": "Li Li\r\nThank you for your suggestion. But I don't quite understand why those data leak feature make the model fail to generalize. Could you explain a little bit?\r\n\r\n@YijunMa\r\nI also try xgboost, I get the similar result as you lol",
    "119638": "Li Li. I am trying to round up the result, hopefully the leakage problem can be solved. The trick is that distance itself might mean something. Say a long distance might means a international travel, and in this case the cluster could narrow to certain range.\r\n\r\nDoing feature engineering in such a big data set is so painful... Do you have any clue now?",
    "119639": "kuan chen.\r\n\r\nWhat I did is:\r\n\r\n 1. 5-fold cv, and all the test result is better than local popular around 0.05 at least.\r\n 2. So excited, and submit my answer (of course updated some of the rows with leakage answers), waiting for top....\r\n 3. Dala, the result is even worse!!\r\n 4. I have no idea why, cuz basically it means cross validation is not working. Is it because of my wrong code? Then I realize, aha, I have another hold out test group maybe I can used for testing, which is the leakage data group. I though since it is part of the LB data, it should be a good test set.\r\n 5. For this test set, the XGB is still better than local popular guess. I am now crazy, and have no idea what to trust...\r\n 6. Until, I check the feature importance (maybe I should do this earlier). the distance feature is almost the only important feature, and the second important one is user's area.\r\n 7. As I wipe these two features, XGB performs worse than local popular. Now that I am almost 80% sure it's a data leakage reason.\r\n\r\nThe idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the \"local top 5 popular hotel model\" perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.",
    "119644": "YijunMa\r\n\r\nThank for share your experiment and result! I did the feature importance first by random forest and get the same observation as yours (but at that time, I didn't know about data leakage would cause severe problem on model..., now I learned). \r\n\r\nI used data leakage solution now but I am still considering whether it is possible to integrate those models and leakage solution",
    "121495": "Hi guys,\r\nI blend my rfr results with leakage and local popular results. I can't pass lb=0.5\r\nwhat's your best lb score getting by using xgb or rfr?",
    "121523": "I am also curious about the best lb score of xgboost (for me the training the process is too slow...)",
    "121528": "With pure XGBoost, my best score is **0.25328**. But that's to weak to use in my submissions at this point. I've tried to mix it in with my counting strategies, but it gets dropped because it's too weak; https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/ .",
    "121644": "YijunMa - why did you wiped the this two features?\r\nnot all of the rows are leaky rows, as you know.",
    "121646": "I didn't exactly wipe it, it just wasn't used by my evolutionary solution; https://lotsacode.wordpress.com/2016/05/26/evolving-a-better-solution/\r\n\r\nThe thing is, these counters don't only work for leaky data. There are public scripts that do >0.5 by using counters. Ie: pick the hotel clusters that best match some set of features in your search. \r\n\r\nThere isn't much Machine Learning involved in hand picking these counters, but as my blog post descibes, I was able to create a system that evolves them.\r\n\r\nThat system had the option to use the XGBoost submission, but it ultimately wasn't used for the best solutions.",
    "121655": "Mattias Fagerlund Thank you for sharing. Did you use any information of the destination file?",
    "121714": "BTW I just ran the .50161 script excluding 2015 training data, and it got .49266.",
    "121726": "happycube  what does that mean? Training data just includes 2013 and 2014 as I remembered.",
    "121728": "Oops - I meant to say year for checkin date (book_year in the script) - of which about 12-13% is dated 2015.\r\n\r\n2013 clicks only gets .40609, 2014 .48472",
    "121910": "\"The idea is that XGB memorize in training group's hotel cluster based on users' location and the distance. And since in the test set, same user location and distance will occur (which correspond to a same hotel cluster), the model perform better than random guess. Notice that the \"local top 5 popular hotel model\" perform worse than XGB only because it does not take the leakage into account, as we update the answer with leakage solution, it is even better than XGB.\"\r\n\r\n@YijunMa  You mean that \"local top 5 popular hotel model\" with leakage solution is better than XGB. Here, is this XGB with leakage solution or without?  Thank you.",
    "121917": "In my tests, I was never able to make XGB perform better than most popular hotels with leakage. I work on a sample of time ordered 10k bookings, and I was barely able to beat pop hotels without leakage, and LB score confirms. It wastes a lot of time for what it performs. As @Mattias, I got better submissions by mixing xgb predictions with different leakage solutions..",
    "121924": "FengLi, my example here is trying to say how leakage data can ruin the XGB model - even worse than local pop5 model. In this case, the XGB model is merged with leakage solution, so that the two models are compare in the same criteria."
  },
  "source": "meta"
}