{
  "id": 21646,
  "title": "Best solution without any widening of the leak",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21646",
  "author_name": "",
  "post_date": "2016-06-13T18:16:30.417Z",
  "votes": 1,
  "comment_count": 10,
  "views": 1628,
  "content": "<p>I'm very surprized  that yet nobody started such a topic.\nCould somebody share the scores obtained without any further leakage exploration (widening the leak, full leak etc.)?\nI will be impressed if somebody could achive 0.52 on LB without any widening of the leak.</p>",
  "messages": [
    {
      "id": "123716",
      "postDate": "06/13/2016 18:16:30",
      "content": "<p>I'm very surprized  that yet nobody started such a topic.\nCould somebody share the scores obtained without any further leakage exploration (widening the leak, full leak etc.)?\nI will be impressed if somebody could achive 0.52 on LB without any widening of the leak.</p>",
      "rawMarkdown": "I'm very surprized  that yet nobody started such a topic.\r\nCould somebody share the scores obtained without any further leakage exploration (widening the leak, full leak etc.)?\r\nI will be impressed if somebody could achive 0.52 on LB without any widening of the leak.",
      "votes": null
    },
    {
      "id": "123719",
      "postDate": "06/13/2016 18:29:05",
      "content": "<p>0.518 - slightly too low to impress you ;)</p>",
      "rawMarkdown": "0.518 - slightly too low to impress you ;)",
      "votes": null
    },
    {
      "id": "123727",
      "postDate": "06/13/2016 18:55:54",
      "content": "<p>[quote=Gert;123719]\n0.518 - slightly too low to impress you ;)\n[/quote]\nI wrote 0.52, but not 0.520. So, 0.518 is sufficient :).</p>",
      "rawMarkdown": "[quote=Gert;123719]\r\n0.518 - slightly too low to impress you ;)\r\n[/quote]\r\nI wrote 0.52, but not 0.520. So, 0.518 is sufficient :).",
      "votes": null
    },
    {
      "id": "123729",
      "postDate": "06/13/2016 19:16:18",
      "content": "<p>0.51583 too? :-)\n(One xgboost model with basic leak in the selected rows)</p>",
      "rawMarkdown": "0.51583 too? :-)\r\n(One xgboost model with basic leak in the selected rows)",
      "votes": null
    },
    {
      "id": "123818",
      "postDate": "06/14/2016 03:35:52",
      "content": "<p>0.38547 VW without leak, without regularization and a lot of parameter space.  </p>",
      "rawMarkdown": "0.38547 VW without leak, without regularization and a lot of parameter space.",
      "votes": null
    },
    {
      "id": "123821",
      "postDate": "06/14/2016 03:46:03",
      "content": "<p>I wounder  how useful results would be for Expedia. Seems like they were interested in sorting hotel clusters for particular user in recommendation page or search. Instead there are smart solution on deanonymization data and figuring out what hotel clusters available for given location, which Expedia already knows.</p>",
      "rawMarkdown": "I wounder  how useful results would be for Expedia. Seems like they were interested in sorting hotel clusters for particular user in recommendation page or search. Instead there are smart solution on deanonymization data and figuring out what hotel clusters available for given location, which Expedia already knows.",
      "votes": null
    },
    {
      "id": "123870",
      "postDate": "06/14/2016 08:42:23",
      "content": "<p>[quote=Jos&#233; A. Guerrero;123729]</p>\n\n<p>0.51583 too? :-)\n(One xgboost model with basic leak in the selected rows)</p>\n\n<p>[/quote]</p>\n\n<p>Hi Jose, are you planning to share your solution?\nVery keen to see how you got there with XGB\nThanks</p>",
      "rawMarkdown": "[quote=José A. Guerrero;123729]\r\n\r\n0.51583 too? :-)\r\n(One xgboost model with basic leak in the selected rows)\r\n\r\n[/quote]\r\n\r\nHi Jose, are you planning to share your solution?\r\nVery keen to see how you got there with XGB\r\nThanks",
      "votes": null
    },
    {
      "id": "123871",
      "postDate": "06/14/2016 08:44:16",
      "content": "<p>@Gert, have you described your method (beyond the widening of the leak), because my XGBoost attempts were pretty bad. I'm be curious what features you used, what depth trees etc!</p>",
      "rawMarkdown": "Gert, have you described your method (beyond the widening of the leak), because my XGBoost attempts were pretty bad. I'm be curious what features you used, what depth trees etc!",
      "votes": null
    },
    {
      "id": "123944",
      "postDate": "06/14/2016 15:41:21",
      "content": "<p>Hi Mattias, some things that helped XGBoost:</p>\n\n<ul>\n<li>use historical booking counts as features (like: number of bookings in same cluster for same search destination)</li>\n<li>train xgboost within a holdout set (like last 4 months) that is not used to create the counts</li>\n<li>expand the train set to 100 cluster-rows for each booking, and then reduce it by using only a subset from the rows where cluster!=bookedcluster (this reduces complexity to 2-class-classification, and also reduces the number of columns - while increasing rows)</li>\n</ul>",
      "rawMarkdown": "Hi Mattias, some things that helped XGBoost:\r\n\r\n - use historical booking counts as features (like: number of bookings in same cluster for same search destination)\r\n - train xgboost within a holdout set (like last 4 months) that is not used to create the counts\r\n - expand the train set to 100 cluster-rows for each booking, and then reduce it by using only a subset from the rows where cluster!=bookedcluster (this reduces complexity to 2-class-classification, and also reduces the number of columns - while increasing rows)",
      "votes": null
    },
    {
      "id": "124054",
      "postDate": "06/15/2016 07:53:59",
      "content": "<p>@Gert, thanks for the elaboration</p>\n\n<h2>use historical bookings counts as features</h2>\n\n<p>I tried that, it may have helped but not enough; i probably didn't use the best counts.</p>\n\n<h2>train xgboost within a holdout set</h2>\n\n<p>Yeah, every time I've tried to use counts of the type you describe, I get over-fitting. I try to remove the specific row I'm evaluating - so that's not included in the count, but I still tend to get over fitting. Using a holdout set seems like a better option.</p>\n\n<h2>expand the train set to 100 cluster-rows for each booking</h2>\n\n<p>very interesting. So for each row you generate 100\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?</p>\n\n<p>/m</p>",
      "rawMarkdown": "Gert, thanks for the elaboration\r\n\r\n## use historical bookings counts as features ##\r\n\r\nI tried that, it may have helped but not enough; i probably didn't use the best counts.\r\n\r\n## train xgboost within a holdout set ##\r\nYeah, every time I've tried to use counts of the type you describe, I get over-fitting. I try to remove the specific row I'm evaluating - so that's not included in the count, but I still tend to get over fitting. Using a holdout set seems like a better option.\r\n\r\n## expand the train set to 100 cluster-rows for each booking ##\r\nvery interesting. So for each row you generate 100\r\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?\r\n\r\n/m",
      "votes": null
    },
    {
      "id": "124077",
      "postDate": "06/15/2016 11:42:03",
      "content": "<p>[quote=Mattias Fagerlund;124054]</p>\n\n<p>So for each row you generate 100\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?</p>\n\n<p>[/quote]</p>\n\n<p>Exactly like that, together with normalized counts (sum to 1.0 over all 100 clusters) for the specific hotel cluster. I haven't tried one forest per class, but I think that a common model has more reliable parameter estimates, especially for the rare classes.</p>",
      "rawMarkdown": "[quote=Mattias Fagerlund;124054]\r\n\r\nSo for each row you generate 100\r\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?\r\n\r\n[/quote]\r\n\r\nExactly like that, together with normalized counts (sum to 1.0 over all 100 clusters) for the specific hotel cluster. I haven't tried one forest per class, but I think that a common model has more reliable parameter estimates, especially for the rare classes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123719,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/13/2016 18:29:05",
      "content": "<p>0.518 - slightly too low to impress you ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123727,
      "author_name": "nedelko",
      "author_url": "",
      "post_date": "06/13/2016 18:55:54",
      "content": "<p>[quote=Gert;123719]\n0.518 - slightly too low to impress you ;)\n[/quote]\nI wrote 0.52, but not 0.520. So, 0.518 is sufficient :).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123729,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/13/2016 19:16:18",
      "content": "<p>0.51583 too? :-)\n(One xgboost model with basic leak in the selected rows)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123818,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "06/14/2016 03:35:52",
      "content": "<p>0.38547 VW without leak, without regularization and a lot of parameter space.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123821,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "06/14/2016 03:46:03",
      "content": "<p>I wounder  how useful results would be for Expedia. Seems like they were interested in sorting hotel clusters for particular user in recommendation page or search. Instead there are smart solution on deanonymization data and figuring out what hotel clusters available for given location, which Expedia already knows.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123870,
      "author_name": "dmenin",
      "author_url": "",
      "post_date": "06/14/2016 08:42:23",
      "content": "<p>[quote=Jos&#233; A. Guerrero;123729]</p>\n\n<p>0.51583 too? :-)\n(One xgboost model with basic leak in the selected rows)</p>\n\n<p>[/quote]</p>\n\n<p>Hi Jose, are you planning to share your solution?\nVery keen to see how you got there with XGB\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123871,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/14/2016 08:44:16",
      "content": "<p>@Gert, have you described your method (beyond the widening of the leak), because my XGBoost attempts were pretty bad. I'm be curious what features you used, what depth trees etc!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123944,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/14/2016 15:41:21",
      "content": "<p>Hi Mattias, some things that helped XGBoost:</p>\n\n<ul>\n<li>use historical booking counts as features (like: number of bookings in same cluster for same search destination)</li>\n<li>train xgboost within a holdout set (like last 4 months) that is not used to create the counts</li>\n<li>expand the train set to 100 cluster-rows for each booking, and then reduce it by using only a subset from the rows where cluster!=bookedcluster (this reduces complexity to 2-class-classification, and also reduces the number of columns - while increasing rows)</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124054,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/15/2016 07:53:59",
      "content": "<p>@Gert, thanks for the elaboration</p>\n\n<h2>use historical bookings counts as features</h2>\n\n<p>I tried that, it may have helped but not enough; i probably didn't use the best counts.</p>\n\n<h2>train xgboost within a holdout set</h2>\n\n<p>Yeah, every time I've tried to use counts of the type you describe, I get over-fitting. I try to remove the specific row I'm evaluating - so that's not included in the count, but I still tend to get over fitting. Using a holdout set seems like a better option.</p>\n\n<h2>expand the train set to 100 cluster-rows for each booking</h2>\n\n<p>very interesting. So for each row you generate 100\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?</p>\n\n<p>/m</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124077,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/15/2016 11:42:03",
      "content": "<p>[quote=Mattias Fagerlund;124054]</p>\n\n<p>So for each row you generate 100\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?</p>\n\n<p>[/quote]</p>\n\n<p>Exactly like that, together with normalized counts (sum to 1.0 over all 100 clusters) for the specific hotel cluster. I haven't tried one forest per class, but I think that a common model has more reliable parameter estimates, especially for the rare classes.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123716": "I'm very surprized  that yet nobody started such a topic.\r\nCould somebody share the scores obtained without any further leakage exploration (widening the leak, full leak etc.)?\r\nI will be impressed if somebody could achive 0.52 on LB without any widening of the leak.",
    "123719": "0.518 - slightly too low to impress you ;)",
    "123727": "[quote=Gert;123719]\r\n0.518 - slightly too low to impress you ;)\r\n[/quote]\r\nI wrote 0.52, but not 0.520. So, 0.518 is sufficient :).",
    "123729": "0.51583 too? :-)\r\n(One xgboost model with basic leak in the selected rows)",
    "123818": "0.38547 VW without leak, without regularization and a lot of parameter space.",
    "123821": "I wounder  how useful results would be for Expedia. Seems like they were interested in sorting hotel clusters for particular user in recommendation page or search. Instead there are smart solution on deanonymization data and figuring out what hotel clusters available for given location, which Expedia already knows.",
    "123870": "[quote=José A. Guerrero;123729]\r\n\r\n0.51583 too? :-)\r\n(One xgboost model with basic leak in the selected rows)\r\n\r\n[/quote]\r\n\r\nHi Jose, are you planning to share your solution?\r\nVery keen to see how you got there with XGB\r\nThanks",
    "123871": "Gert, have you described your method (beyond the widening of the leak), because my XGBoost attempts were pretty bad. I'm be curious what features you used, what depth trees etc!",
    "123944": "Hi Mattias, some things that helped XGBoost:\r\n\r\n - use historical booking counts as features (like: number of bookings in same cluster for same search destination)\r\n - train xgboost within a holdout set (like last 4 months) that is not used to create the counts\r\n - expand the train set to 100 cluster-rows for each booking, and then reduce it by using only a subset from the rows where cluster!=bookedcluster (this reduces complexity to 2-class-classification, and also reduces the number of columns - while increasing rows)",
    "124054": "Gert, thanks for the elaboration\r\n\r\n## use historical bookings counts as features ##\r\n\r\nI tried that, it may have helped but not enough; i probably didn't use the best counts.\r\n\r\n## train xgboost within a holdout set ##\r\nYeah, every time I've tried to use counts of the type you describe, I get over-fitting. I try to remove the specific row I'm evaluating - so that's not included in the count, but I still tend to get over fitting. Using a holdout set seems like a better option.\r\n\r\n## expand the train set to 100 cluster-rows for each booking ##\r\nvery interesting. So for each row you generate 100\r\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?\r\n\r\n/m",
    "124077": "[quote=Mattias Fagerlund;124054]\r\n\r\nSo for each row you generate 100\r\n rows that contain the hotel_cluster as a feature and another feature stating if the cluster was incorrect (99 rows) or correct (1 row)? Did that improve quality over the standard method of creating one forest per class?\r\n\r\n[/quote]\r\n\r\nExactly like that, together with normalized counts (sum to 1.0 over all 100 clusters) for the specific hotel cluster. I haven't tried one forest per class, but I think that a common model has more reliable parameter estimates, especially for the rare classes."
  },
  "source": "meta"
}