{
  "id": 21066,
  "title": "Difference between local validation score and LB score",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21066",
  "author_name": "",
  "post_date": "2016-05-19T15:32:44.547Z",
  "votes": 1,
  "comment_count": 10,
  "views": 1183,
  "content": "<p>Hi all,</p>\n\n<p>I am seeing a huge difference between my local validation score and my LB score. \nBasically, I am getting data for 20000 users (around 250k rows), doing data clean-up and manipulation and then splinting it into train (80%) and validation(20%) using StratifiedShuffleSplit on hotel_cluster. (I keep both is_booking = 0 and 1 on my train and validation).\nThen I fit a randon forest model and my MAP@5 is around 0.35\nWhen I apply my model to the test data, my leaderbord score is 0.21.</p>\n\n<p>What do you guys think I'm doing wrong to justify such a difference?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "120624",
      "postDate": "05/19/2016 15:32:44",
      "content": "<p>Hi all,</p>\n\n<p>I am seeing a huge difference between my local validation score and my LB score. \nBasically, I am getting data for 20000 users (around 250k rows), doing data clean-up and manipulation and then splinting it into train (80%) and validation(20%) using StratifiedShuffleSplit on hotel_cluster. (I keep both is_booking = 0 and 1 on my train and validation).\nThen I fit a randon forest model and my MAP@5 is around 0.35\nWhen I apply my model to the test data, my leaderbord score is 0.21.</p>\n\n<p>What do you guys think I'm doing wrong to justify such a difference?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi all,\r\n\r\nI am seeing a huge difference between my local validation score and my LB score. \r\nBasically, I am getting data for 20000 users (around 250k rows), doing data clean-up and manipulation and then splinting it into train (80%) and validation(20%) using StratifiedShuffleSplit on hotel_cluster. (I keep both is_booking = 0 and 1 on my train and validation).\r\nThen I fit a randon forest model and my MAP@5 is around 0.35\r\nWhen I apply my model to the test data, my leaderbord score is 0.21.\r\n\r\nWhat do you guys think I'm doing wrong to justify such a difference?\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "120709",
      "postDate": "05/20/2016 02:14:56",
      "content": "<p>That's strange! Have you tried keeping only bookings in validation? Because the test data is only bookings...</p>",
      "rawMarkdown": "That's strange! Have you tried keeping only bookings in validation? Because the test data is only bookings...",
      "votes": null
    },
    {
      "id": "120982",
      "postDate": "05/22/2016 12:05:55",
      "content": "<p>did you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)</p>",
      "rawMarkdown": "did you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)",
      "votes": null
    },
    {
      "id": "120991",
      "postDate": "05/22/2016 13:29:51",
      "content": "<p>definitely the &quot;is_booking&quot; flag.\nI had the same issue till I removed the is_booking flag.</p>\n\n<p>You could two assign some weight to the is_booking = 0 samples in order to learn from them instead of just drop them.</p>",
      "rawMarkdown": "definitely the \"is_booking\" flag.\r\nI had the same issue till I removed the is_booking flag.\r\n\r\nYou could two assign some weight to the is_booking = 0 samples in order to learn from them instead of just drop them.",
      "votes": null
    },
    {
      "id": "120993",
      "postDate": "05/22/2016 13:54:59",
      "content": "<p>@ Arnold Zephir</p>\n\n<p>Hi I got the same problem by using random forest, For assign the weight to click event (is_booking==0), do you mean that I can assign o.15 to is_booking==0 and 0.85 to is_booking==1  ?</p>",
      "rawMarkdown": "Arnold Zephir\r\n\r\nHi I got the same problem by using random forest, For assign the weight to click event (is_booking==0), do you mean that I can assign o.15 to is_booking==0 and 0.85 to is_booking==1  ?",
      "votes": null
    },
    {
      "id": "120999",
      "postDate": "05/22/2016 15:58:13",
      "content": "<p>yes I do mean that.\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful  </p>",
      "rawMarkdown": "yes I do mean that.\r\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful",
      "votes": null
    },
    {
      "id": "121002",
      "postDate": "05/22/2016 16:26:21",
      "content": "<p>@Arnold Zephir</p>\n\n<p>Thanks a lot! May I ask how long do you train your xgboost model? I tried it before but it took long time, and the score is not good</p>",
      "rawMarkdown": "Arnold Zephir\r\n\r\nThanks a lot! May I ask how long do you train your xgboost model? I tried it before but it took long time, and the score is not good",
      "votes": null
    },
    {
      "id": "121005",
      "postDate": "05/22/2016 17:27:53",
      "content": "<p>It depends a lot ( of course ) on two parameters : \nmax_depth and nb_round</p>\n\n<p>with an 8 max_depth, I got about 30s per round, time 200 round (so 100mn )\nI got a octocore with 32 ram (thus, no swap )</p>\n\n<p>Yet, XGBoost is not my prime model on this kaggle, it does not perform very well du to the lot of one hot encoded features.</p>\n\n<p>Tree are not the best models to my opinion.\nit' too long for iterate</p>",
      "rawMarkdown": "It depends a lot ( of course ) on two parameters : \r\nmax_depth and nb_round\r\n\r\nwith an 8 max_depth, I got about 30s per round, time 200 round (so 100mn )\r\nI got a octocore with 32 ram (thus, no swap )\r\n\r\nYet, XGBoost is not my prime model on this kaggle, it does not perform very well du to the lot of one hot encoded features.\r\n\r\nTree are not the best models to my opinion.\r\nit' too long for iterate",
      "votes": null
    },
    {
      "id": "121057",
      "postDate": "05/23/2016 10:31:17",
      "content": "<p>[quote=KnitCode;120982]</p>\n\n<p>did you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)</p>\n\n<p>[/quote]</p>\n\n<p>hey, Im splititng the data like so:</p>\n\n<blockquote>\n  <p>StratifiedShuffleSplit(X['hotel_cluster'].astype(int), 1,\n  test_size=0.2, random_state=123456)</p>\n</blockquote>\n\n<p>so in the end are you using the date_time split? It does produce a similar result to the LB, Im just wondering why the StratifiedShuffleSplit produces such a  better validation score thats not reflected on the LB</p>\n\n<p>thanks</p>",
      "rawMarkdown": "[quote=KnitCode;120982]\r\n\r\ndid you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)\r\n\r\n[/quote]\r\n\r\nhey, Im splititng the data like so:\r\n\r\n> StratifiedShuffleSplit(X['hotel_cluster'].astype(int), 1,\r\n> test_size=0.2, random_state=123456)\r\n\r\n\r\nso in the end are you using the date_time split? It does produce a similar result to the LB, Im just wondering why the StratifiedShuffleSplit produces such a  better validation score thats not reflected on the LB\r\n\r\nthanks",
      "votes": null
    },
    {
      "id": "121058",
      "postDate": "05/23/2016 10:32:33",
      "content": "<p>[quote=Arnold Zephir;120999]</p>\n\n<p>yes I do mean that.\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful  </p>\n\n<p>[/quote]</p>\n\n<p>hi, can you show me how to do that?\nthanks</p>",
      "rawMarkdown": "[quote=Arnold Zephir;120999]\r\n\r\nyes I do mean that.\r\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful  \r\n\r\n[/quote]\r\n\r\n\r\nhi, can you show me how to do that?\r\nthanks",
      "votes": null
    },
    {
      "id": "121063",
      "postDate": "05/23/2016 11:48:33",
      "content": "<p>@italo, </p>\n\n<p>i am by no means an expert, but i see a few sources of potential problems with the stratified sample and keeping the data the way you are doing so. </p>\n\n<p>first, as folks mentioned, if you use click only (booking=0) data to validate from, you're validating on things that are not representative of the LB test set. most clicks don't convert to bookings, and yet they represent the vast majority of the data, so however you built your model, you'd likely have it biased toward classifying clicks, not books... whereas the test test set is on bookings. So, definitely want to remove the booking=0 from your local test/validation set.  </p>\n\n<p>When you do that, you'll find that the number of unique users you have in the validation set is a lot smaller than 20,000.  i think the impact of that would vary depending on your specific model approach.  i'm guessing you'll have closer to 5-6k unique users.  also the validation set will be greatly reduced. </p>\n\n<p>Second, as relates to the date-time split... i think the issue here is that you are bleeding information from the training to the validation set. the stratified split is ensure that the hotel clusters are evenly represented in the training and validation set, but because date-time isn't a factor, you are very likely to have your validation items as part of a longer chain of events that are in the training data. while you haven't duplicated those individual events, you've duplicated a huge amount of the information in them. </p>\n\n<p>you can see this just by manually eying the data. for any given booking event in your validation/test set, there is a very high likelihood that all the other events nearby in time are in your training set. and those events are likely to contain the truth in some way, e.g., they clicked into that cluster before they booked it. </p>\n\n<p>on the other hand, the LB is based on data in 2015 for which there is no user history nearby in time. putting aside the leak between sets that is well documented, the training data doesn't match the exact search events that led to the 2015 booking, meaning that they are different trips.  this is the scenario you want to reproduce. so, if you first take your 250k rows (or more), split them by date time, and then take your validation as the latter period bookings, you'll have a better match.  dropping a few months in the middle might even be better simulation. </p>\n\n<p>you're lucky you only dropped from 0.35 to 0.21, mine dropped from 0.35 to 0.08. :)</p>",
      "rawMarkdown": "italo, \r\n\r\ni am by no means an expert, but i see a few sources of potential problems with the stratified sample and keeping the data the way you are doing so. \r\n\r\nfirst, as folks mentioned, if you use click only (booking=0) data to validate from, you're validating on things that are not representative of the LB test set. most clicks don't convert to bookings, and yet they represent the vast majority of the data, so however you built your model, you'd likely have it biased toward classifying clicks, not books... whereas the test test set is on bookings. So, definitely want to remove the booking=0 from your local test/validation set.  \r\n\r\nWhen you do that, you'll find that the number of unique users you have in the validation set is a lot smaller than 20,000.  i think the impact of that would vary depending on your specific model approach.  i'm guessing you'll have closer to 5-6k unique users.  also the validation set will be greatly reduced. \r\n\r\nSecond, as relates to the date-time split... i think the issue here is that you are bleeding information from the training to the validation set. the stratified split is ensure that the hotel clusters are evenly represented in the training and validation set, but because date-time isn't a factor, you are very likely to have your validation items as part of a longer chain of events that are in the training data. while you haven't duplicated those individual events, you've duplicated a huge amount of the information in them. \r\n\r\nyou can see this just by manually eying the data. for any given booking event in your validation/test set, there is a very high likelihood that all the other events nearby in time are in your training set. and those events are likely to contain the truth in some way, e.g., they clicked into that cluster before they booked it. \r\n\r\non the other hand, the LB is based on data in 2015 for which there is no user history nearby in time. putting aside the leak between sets that is well documented, the training data doesn't match the exact search events that led to the 2015 booking, meaning that they are different trips.  this is the scenario you want to reproduce. so, if you first take your 250k rows (or more), split them by date time, and then take your validation as the latter period bookings, you'll have a better match.  dropping a few months in the middle might even be better simulation. \r\n\r\nyou're lucky you only dropped from 0.35 to 0.21, mine dropped from 0.35 to 0.08. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120709,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "05/20/2016 02:14:56",
      "content": "<p>That's strange! Have you tried keeping only bookings in validation? Because the test data is only bookings...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120982,
      "author_name": "knitcode",
      "author_url": "",
      "post_date": "05/22/2016 12:05:55",
      "content": "<p>did you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120991,
      "author_name": "zarnold",
      "author_url": "",
      "post_date": "05/22/2016 13:29:51",
      "content": "<p>definitely the &quot;is_booking&quot; flag.\nI had the same issue till I removed the is_booking flag.</p>\n\n<p>You could two assign some weight to the is_booking = 0 samples in order to learn from them instead of just drop them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120993,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/22/2016 13:54:59",
      "content": "<p>@ Arnold Zephir</p>\n\n<p>Hi I got the same problem by using random forest, For assign the weight to click event (is_booking==0), do you mean that I can assign o.15 to is_booking==0 and 0.85 to is_booking==1  ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120999,
      "author_name": "zarnold",
      "author_url": "",
      "post_date": "05/22/2016 15:58:13",
      "content": "<p>yes I do mean that.\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121002,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/22/2016 16:26:21",
      "content": "<p>@Arnold Zephir</p>\n\n<p>Thanks a lot! May I ask how long do you train your xgboost model? I tried it before but it took long time, and the score is not good</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121005,
      "author_name": "zarnold",
      "author_url": "",
      "post_date": "05/22/2016 17:27:53",
      "content": "<p>It depends a lot ( of course ) on two parameters : \nmax_depth and nb_round</p>\n\n<p>with an 8 max_depth, I got about 30s per round, time 200 round (so 100mn )\nI got a octocore with 32 ram (thus, no swap )</p>\n\n<p>Yet, XGBoost is not my prime model on this kaggle, it does not perform very well du to the lot of one hot encoded features.</p>\n\n<p>Tree are not the best models to my opinion.\nit' too long for iterate</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121057,
      "author_name": "italotf",
      "author_url": "",
      "post_date": "05/23/2016 10:31:17",
      "content": "<p>[quote=KnitCode;120982]</p>\n\n<p>did you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)</p>\n\n<p>[/quote]</p>\n\n<p>hey, Im splititng the data like so:</p>\n\n<blockquote>\n  <p>StratifiedShuffleSplit(X['hotel_cluster'].astype(int), 1,\n  test_size=0.2, random_state=123456)</p>\n</blockquote>\n\n<p>so in the end are you using the date_time split? It does produce a similar result to the LB, Im just wondering why the StratifiedShuffleSplit produces such a  better validation score thats not reflected on the LB</p>\n\n<p>thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121058,
      "author_name": "italotf",
      "author_url": "",
      "post_date": "05/23/2016 10:32:33",
      "content": "<p>[quote=Arnold Zephir;120999]</p>\n\n<p>yes I do mean that.\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful  </p>\n\n<p>[/quote]</p>\n\n<p>hi, can you show me how to do that?\nthanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121063,
      "author_name": "knitcode",
      "author_url": "",
      "post_date": "05/23/2016 11:48:33",
      "content": "<p>@italo, </p>\n\n<p>i am by no means an expert, but i see a few sources of potential problems with the stratified sample and keeping the data the way you are doing so. </p>\n\n<p>first, as folks mentioned, if you use click only (booking=0) data to validate from, you're validating on things that are not representative of the LB test set. most clicks don't convert to bookings, and yet they represent the vast majority of the data, so however you built your model, you'd likely have it biased toward classifying clicks, not books... whereas the test test set is on bookings. So, definitely want to remove the booking=0 from your local test/validation set.  </p>\n\n<p>When you do that, you'll find that the number of unique users you have in the validation set is a lot smaller than 20,000.  i think the impact of that would vary depending on your specific model approach.  i'm guessing you'll have closer to 5-6k unique users.  also the validation set will be greatly reduced. </p>\n\n<p>Second, as relates to the date-time split... i think the issue here is that you are bleeding information from the training to the validation set. the stratified split is ensure that the hotel clusters are evenly represented in the training and validation set, but because date-time isn't a factor, you are very likely to have your validation items as part of a longer chain of events that are in the training data. while you haven't duplicated those individual events, you've duplicated a huge amount of the information in them. </p>\n\n<p>you can see this just by manually eying the data. for any given booking event in your validation/test set, there is a very high likelihood that all the other events nearby in time are in your training set. and those events are likely to contain the truth in some way, e.g., they clicked into that cluster before they booked it. </p>\n\n<p>on the other hand, the LB is based on data in 2015 for which there is no user history nearby in time. putting aside the leak between sets that is well documented, the training data doesn't match the exact search events that led to the 2015 booking, meaning that they are different trips.  this is the scenario you want to reproduce. so, if you first take your 250k rows (or more), split them by date time, and then take your validation as the latter period bookings, you'll have a better match.  dropping a few months in the middle might even be better simulation. </p>\n\n<p>you're lucky you only dropped from 0.35 to 0.21, mine dropped from 0.35 to 0.08. :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120624": "Hi all,\r\n\r\nI am seeing a huge difference between my local validation score and my LB score. \r\nBasically, I am getting data for 20000 users (around 250k rows), doing data clean-up and manipulation and then splinting it into train (80%) and validation(20%) using StratifiedShuffleSplit on hotel_cluster. (I keep both is_booking = 0 and 1 on my train and validation).\r\nThen I fit a randon forest model and my MAP@5 is around 0.35\r\nWhen I apply my model to the test data, my leaderbord score is 0.21.\r\n\r\nWhat do you guys think I'm doing wrong to justify such a difference?\r\n\r\nThanks",
    "120709": "That's strange! Have you tried keeping only bookings in validation? Because the test data is only bookings...",
    "120982": "did you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)",
    "120991": "definitely the \"is_booking\" flag.\r\nI had the same issue till I removed the is_booking flag.\r\n\r\nYou could two assign some weight to the is_booking = 0 samples in order to learn from them instead of just drop them.",
    "120993": "Arnold Zephir\r\n\r\nHi I got the same problem by using random forest, For assign the weight to click event (is_booking==0), do you mean that I can assign o.15 to is_booking==0 and 0.85 to is_booking==1  ?",
    "120999": "yes I do mean that.\r\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful",
    "121002": "Arnold Zephir\r\n\r\nThanks a lot! May I ask how long do you train your xgboost model? I tried it before but it took long time, and the score is not good",
    "121005": "It depends a lot ( of course ) on two parameters : \r\nmax_depth and nb_round\r\n\r\nwith an 8 max_depth, I got about 30s per round, time 200 round (so 100mn )\r\nI got a octocore with 32 ram (thus, no swap )\r\n\r\nYet, XGBoost is not my prime model on this kaggle, it does not perform very well du to the lot of one hot encoded features.\r\n\r\nTree are not the best models to my opinion.\r\nit' too long for iterate",
    "121057": "[quote=KnitCode;120982]\r\n\r\ndid you split your train and validation by date_time, or random sampling? I found a large drop in performance using a date_time split vs random sampling.  I had found the same issue you saw in my local validation vs the leaderboard, and so i went back and hunted it down.  in hindsight, somewhat obvious, but at the time random sampling seemed fine. :)\r\n\r\n[/quote]\r\n\r\nhey, Im splititng the data like so:\r\n\r\n> StratifiedShuffleSplit(X['hotel_cluster'].astype(int), 1,\r\n> test_size=0.2, random_state=123456)\r\n\r\n\r\nso in the end are you using the date_time split? It does produce a similar result to the LB, Im just wondering why the StratifiedShuffleSplit produces such a  better validation score thats not reflected on the LB\r\n\r\nthanks",
    "121058": "[quote=Arnold Zephir;120999]\r\n\r\nyes I do mean that.\r\nIn XGBoost for example, it increases the weight for gini criterion. It's sometimes useful  \r\n\r\n[/quote]\r\n\r\n\r\nhi, can you show me how to do that?\r\nthanks",
    "121063": "italo, \r\n\r\ni am by no means an expert, but i see a few sources of potential problems with the stratified sample and keeping the data the way you are doing so. \r\n\r\nfirst, as folks mentioned, if you use click only (booking=0) data to validate from, you're validating on things that are not representative of the LB test set. most clicks don't convert to bookings, and yet they represent the vast majority of the data, so however you built your model, you'd likely have it biased toward classifying clicks, not books... whereas the test test set is on bookings. So, definitely want to remove the booking=0 from your local test/validation set.  \r\n\r\nWhen you do that, you'll find that the number of unique users you have in the validation set is a lot smaller than 20,000.  i think the impact of that would vary depending on your specific model approach.  i'm guessing you'll have closer to 5-6k unique users.  also the validation set will be greatly reduced. \r\n\r\nSecond, as relates to the date-time split... i think the issue here is that you are bleeding information from the training to the validation set. the stratified split is ensure that the hotel clusters are evenly represented in the training and validation set, but because date-time isn't a factor, you are very likely to have your validation items as part of a longer chain of events that are in the training data. while you haven't duplicated those individual events, you've duplicated a huge amount of the information in them. \r\n\r\nyou can see this just by manually eying the data. for any given booking event in your validation/test set, there is a very high likelihood that all the other events nearby in time are in your training set. and those events are likely to contain the truth in some way, e.g., they clicked into that cluster before they booked it. \r\n\r\non the other hand, the LB is based on data in 2015 for which there is no user history nearby in time. putting aside the leak between sets that is well documented, the training data doesn't match the exact search events that led to the 2015 booking, meaning that they are different trips.  this is the scenario you want to reproduce. so, if you first take your 250k rows (or more), split them by date time, and then take your validation as the latter period bookings, you'll have a better match.  dropping a few months in the middle might even be better simulation. \r\n\r\nyou're lucky you only dropped from 0.35 to 0.21, mine dropped from 0.35 to 0.08. :)"
  },
  "source": "meta"
}