{
  "id": 20737,
  "title": "Validation vs LB",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20737",
  "author_name": "",
  "post_date": "2016-05-05T15:26:19.050Z",
  "votes": 21,
  "comment_count": 9,
  "views": 2068,
  "content": "<p>If anybody cares here is my relationship between local validation (trained on 2013, validated on 2014 (or vice versa ?)) and LB score (trained on all data)</p>",
  "messages": [
    {
      "id": "118829",
      "postDate": "05/05/2016 15:26:19",
      "content": "<p>If anybody cares here is my relationship between local validation (trained on 2013, validated on 2014 (or vice versa ?)) and LB score (trained on all data)</p>",
      "rawMarkdown": "If anybody cares here is my relationship between local validation (trained on 2013, validated on 2014 (or vice versa ?)) and LB score (trained on all data)",
      "votes": null
    },
    {
      "id": "119589",
      "postDate": "05/11/2016 16:02:08",
      "content": "<p>Very interesting and useful information! Just to confirm, you did test on all of 2014 data ?</p>\n\n<p>Also one more question: since there are a lot more data in 2014 than 2013, did you (or anyone) tried the opposite ? (training on 2014 and testing on 2013).\nI would be curious to see how that correlate with LB.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Very interesting and useful information! Just to confirm, you did test on all of 2014 data ?\r\n\r\nAlso one more question: since there are a lot more data in 2014 than 2013, did you (or anyone) tried the opposite ? (training on 2014 and testing on 2013).\r\nI would be curious to see how that correlate with LB.\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "119594",
      "postDate": "05/11/2016 17:11:18",
      "content": "<p>It is funny. I checked my code and found out that this is exactly what I did (train on 2014, validate on 2013) :) </p>",
      "rawMarkdown": "It is funny. I checked my code and found out that this is exactly what I did (train on 2014, validate on 2013) :)",
      "votes": null
    },
    {
      "id": "119598",
      "postDate": "05/11/2016 17:29:55",
      "content": "<p>Great ! That explain why I couldn't match your MAP :) </p>",
      "rawMarkdown": "Great ! That explain why I couldn't match your MAP :)",
      "votes": null
    },
    {
      "id": "119716",
      "postDate": "05/12/2016 13:19:44",
      "content": "<p>Great!, one more question, does your 2013 dataset only contain is_booking ==1 or it contains both is_booking==1 and 0? </p>",
      "rawMarkdown": "Great!, one more question, does your 2013 dataset only contain is_booking ==1 or it contains both is_booking==1 and 0?",
      "votes": null
    },
    {
      "id": "119792",
      "postDate": "05/12/2016 20:00:25",
      "content": "<p>Training dataset contains both, validation dataset contains only is_booking==1</p>",
      "rawMarkdown": "Training dataset contains both, validation dataset contains only is_booking==1",
      "votes": null
    },
    {
      "id": "119959",
      "postDate": "05/14/2016 03:29:31",
      "content": "<p>Maybe also test on the leakage data set - 2015 data? Mine also drop about 0.03 when train in 2014 data and 2015 leakage data. I guess it's a yearly difference, and pretty hard to avoid if only using 2014 data to train the model.</p>",
      "rawMarkdown": "Maybe also test on the leakage data set - 2015 data? Mine also drop about 0.03 when train in 2014 data and 2015 leakage data. I guess it's a yearly difference, and pretty hard to avoid if only using 2014 data to train the model.",
      "votes": null
    },
    {
      "id": "121752",
      "postDate": "05/29/2016 12:32:29",
      "content": "<p>@Sergey Yurgenson - This is very useful.\nHowever in order to check ideas, training all entire 2014, and validating on 2013 is slow.\nAre you using some sample out of this,  to test your ideas?\nAnd if so, which part do you use for test and which for validation?\nThank you. </p>",
      "rawMarkdown": "Sergey Yurgenson - This is very useful.\r\nHowever in order to check ideas, training all entire 2014, and validating on 2013 is slow.\r\nAre you using some sample out of this,  to test your ideas?\r\nAnd if so, which part do you use for test and which for validation?\r\nThank you.",
      "votes": null
    },
    {
      "id": "121765",
      "postDate": "05/29/2016 15:10:31",
      "content": "<p>@Sergey Yurgenson - Thanks. What class of ML algorithm you used to train your model?</p>",
      "rawMarkdown": "Sergey Yurgenson - Thanks. What class of ML algorithm you used to train your model?",
      "votes": null
    },
    {
      "id": "121786",
      "postDate": "05/29/2016 18:13:56",
      "content": "<p>Interesting work. For me this relationship looks a lot different, but I may be calculating my validation accuracy wrong. I just take the percentage of cases where the correct cluster is in the predicted top 5. Is this more or less how map5 is calculated?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Interesting work. For me this relationship looks a lot different, but I may be calculating my validation accuracy wrong. I just take the percentage of cases where the correct cluster is in the predicted top 5. Is this more or less how map5 is calculated?\r\n\r\nThanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119589,
      "author_name": "romains",
      "author_url": "",
      "post_date": "05/11/2016 16:02:08",
      "content": "<p>Very interesting and useful information! Just to confirm, you did test on all of 2014 data ?</p>\n\n<p>Also one more question: since there are a lot more data in 2014 than 2013, did you (or anyone) tried the opposite ? (training on 2014 and testing on 2013).\nI would be curious to see how that correlate with LB.</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119594,
      "author_name": "ccccat",
      "author_url": "",
      "post_date": "05/11/2016 17:11:18",
      "content": "<p>It is funny. I checked my code and found out that this is exactly what I did (train on 2014, validate on 2013) :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119598,
      "author_name": "romains",
      "author_url": "",
      "post_date": "05/11/2016 17:29:55",
      "content": "<p>Great ! That explain why I couldn't match your MAP :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119716,
      "author_name": "kuanchen",
      "author_url": "",
      "post_date": "05/12/2016 13:19:44",
      "content": "<p>Great!, one more question, does your 2013 dataset only contain is_booking ==1 or it contains both is_booking==1 and 0? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119792,
      "author_name": "ccccat",
      "author_url": "",
      "post_date": "05/12/2016 20:00:25",
      "content": "<p>Training dataset contains both, validation dataset contains only is_booking==1</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119959,
      "author_name": "ma350365879",
      "author_url": "",
      "post_date": "05/14/2016 03:29:31",
      "content": "<p>Maybe also test on the leakage data set - 2015 data? Mine also drop about 0.03 when train in 2014 data and 2015 leakage data. I guess it's a yearly difference, and pretty hard to avoid if only using 2014 data to train the model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121752,
      "author_name": "dot277",
      "author_url": "",
      "post_date": "05/29/2016 12:32:29",
      "content": "<p>@Sergey Yurgenson - This is very useful.\nHowever in order to check ideas, training all entire 2014, and validating on 2013 is slow.\nAre you using some sample out of this,  to test your ideas?\nAnd if so, which part do you use for test and which for validation?\nThank you. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121765,
      "author_name": "zyazzy",
      "author_url": "",
      "post_date": "05/29/2016 15:10:31",
      "content": "<p>@Sergey Yurgenson - Thanks. What class of ML algorithm you used to train your model?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121786,
      "author_name": "sanderdalm",
      "author_url": "",
      "post_date": "05/29/2016 18:13:56",
      "content": "<p>Interesting work. For me this relationship looks a lot different, but I may be calculating my validation accuracy wrong. I just take the percentage of cases where the correct cluster is in the predicted top 5. Is this more or less how map5 is calculated?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118829": "If anybody cares here is my relationship between local validation (trained on 2013, validated on 2014 (or vice versa ?)) and LB score (trained on all data)",
    "119589": "Very interesting and useful information! Just to confirm, you did test on all of 2014 data ?\r\n\r\nAlso one more question: since there are a lot more data in 2014 than 2013, did you (or anyone) tried the opposite ? (training on 2014 and testing on 2013).\r\nI would be curious to see how that correlate with LB.\r\n\r\nThanks",
    "119594": "It is funny. I checked my code and found out that this is exactly what I did (train on 2014, validate on 2013) :)",
    "119598": "Great ! That explain why I couldn't match your MAP :)",
    "119716": "Great!, one more question, does your 2013 dataset only contain is_booking ==1 or it contains both is_booking==1 and 0?",
    "119792": "Training dataset contains both, validation dataset contains only is_booking==1",
    "119959": "Maybe also test on the leakage data set - 2015 data? Mine also drop about 0.03 when train in 2014 data and 2015 leakage data. I guess it's a yearly difference, and pretty hard to avoid if only using 2014 data to train the model.",
    "121752": "Sergey Yurgenson - This is very useful.\r\nHowever in order to check ideas, training all entire 2014, and validating on 2013 is slow.\r\nAre you using some sample out of this,  to test your ideas?\r\nAnd if so, which part do you use for test and which for validation?\r\nThank you.",
    "121765": "Sergey Yurgenson - Thanks. What class of ML algorithm you used to train your model?",
    "121786": "Interesting work. For me this relationship looks a lot different, but I may be calculating my validation accuracy wrong. I just take the percentage of cases where the correct cluster is in the predicted top 5. Is this more or less how map5 is calculated?\r\n\r\nThanks"
  },
  "source": "meta"
}