{
  "id": 20853,
  "title": "Update from the machine room",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20853",
  "author_name": "",
  "post_date": "2016-05-11T08:06:28.973Z",
  "votes": 4,
  "comment_count": 4,
  "views": 1033,
  "content": "<p>I learned so much from the Kaggle community that I wanted to share some information for this competition. Not too specific and not a script at the moment of course, but still.</p>\n\n<p>I did not yet reach 0.5 but only 0.00022 is missing...</p>\n\n<p>Fist of all I created a train and a validation partition randomly (no time split!) using only 10% of the train data. I thought of creating further such sets to do cross validation but so far the single set served me well.</p>\n\n<p>This needs around 3 minutes for a validation and evaluation run in my perl script per try. This allows for a lot of trial and error when looking for the right parameters.</p>\n\n<p>The script takes 40 minutes at the moment for the complete data set. I run it usually 2-3 times per day, depending on the results of the validation runs.</p>\n\n<p>In the script I use some of the data leak solutions but I also created some features/cluster sets on my own. This results in several sets of probable hotel clusters from which I can choose. Choosing and sorting of course is the crucial part.</p>\n\n<p>When doing a validation run I look at the results and how they were achieved. I evaluate hits and misses.</p>\n\n<p>I found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction! When looking at the misses I could mostly not see how I could have correctly forcasted them. So I think the rest of the work is either create new sets of possible clusters or rearrange the current predictions somehow.</p>\n\n<p>By the way: I created another perl script that can merge submission files in a weighted way. So if you have two or more good submission files you possibly get a better one afterwards.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
  "messages": [
    {
      "id": "119533",
      "postDate": "05/11/2016 08:06:28",
      "content": "<p>I learned so much from the Kaggle community that I wanted to share some information for this competition. Not too specific and not a script at the moment of course, but still.</p>\n\n<p>I did not yet reach 0.5 but only 0.00022 is missing...</p>\n\n<p>Fist of all I created a train and a validation partition randomly (no time split!) using only 10% of the train data. I thought of creating further such sets to do cross validation but so far the single set served me well.</p>\n\n<p>This needs around 3 minutes for a validation and evaluation run in my perl script per try. This allows for a lot of trial and error when looking for the right parameters.</p>\n\n<p>The script takes 40 minutes at the moment for the complete data set. I run it usually 2-3 times per day, depending on the results of the validation runs.</p>\n\n<p>In the script I use some of the data leak solutions but I also created some features/cluster sets on my own. This results in several sets of probable hotel clusters from which I can choose. Choosing and sorting of course is the crucial part.</p>\n\n<p>When doing a validation run I look at the results and how they were achieved. I evaluate hits and misses.</p>\n\n<p>I found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction! When looking at the misses I could mostly not see how I could have correctly forcasted them. So I think the rest of the work is either create new sets of possible clusters or rearrange the current predictions somehow.</p>\n\n<p>By the way: I created another perl script that can merge submission files in a weighted way. So if you have two or more good submission files you possibly get a better one afterwards.</p>\n\n<p>Cheers</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "I learned so much from the Kaggle community that I wanted to share some information for this competition. Not too specific and not a script at the moment of course, but still.\r\n\r\nI did not yet reach 0.5 but only 0.00022 is missing...\r\n\r\nFist of all I created a train and a validation partition randomly (no time split!) using only 10% of the train data. I thought of creating further such sets to do cross validation but so far the single set served me well.\r\n\r\nThis needs around 3 minutes for a validation and evaluation run in my perl script per try. This allows for a lot of trial and error when looking for the right parameters.\r\n\r\nThe script takes 40 minutes at the moment for the complete data set. I run it usually 2-3 times per day, depending on the results of the validation runs.\r\n\r\nIn the script I use some of the data leak solutions but I also created some features/cluster sets on my own. This results in several sets of probable hotel clusters from which I can choose. Choosing and sorting of course is the crucial part.\r\n\r\nWhen doing a validation run I look at the results and how they were achieved. I evaluate hits and misses.\r\n\r\nI found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction! When looking at the misses I could mostly not see how I could have correctly forcasted them. So I think the rest of the work is either create new sets of possible clusters or rearrange the current predictions somehow.\r\n\r\nBy the way: I created another perl script that can merge submission files in a weighted way. So if you have two or more good submission files you possibly get a better one afterwards.\r\n\r\nCheers\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "119578",
      "postDate": "05/11/2016 14:07:40",
      "content": "<p>Thanks for the insights!</p>\n\n<p>I have a similar pipeline, except that I'm not getting such a good score, yet :)</p>\n\n<p>Did you only keep record with is_booking = 1 when creating your validation set ?</p>",
      "rawMarkdown": "Thanks for the insights!\r\n\r\nI have a similar pipeline, except that I'm not getting such a good score, yet :)\r\n\r\nDid you only keep record with is_booking = 1 when creating your validation set ?",
      "votes": null
    },
    {
      "id": "119580",
      "postDate": "05/11/2016 14:38:15",
      "content": "<p>Yes, I selected 30.000 instances from the 2 year period where is_booking == 1. I did this because someone if not the Admin hinted that there is seasonality involved.</p>",
      "rawMarkdown": "Yes, I selected 30.000 instances from the 2 year period where is_booking == 1. I did this because someone if not the Admin hinted that there is seasonality involved.",
      "votes": null
    },
    {
      "id": "119676",
      "postDate": "05/12/2016 07:09:38",
      "content": "<p>[quote=MightyBird;119533]\nI found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction!\n[/quote]</p>\n\n<p>@MightyBird, I can't believe. I have divided the data set in two parts: one with data leak and second one containing the rest. My local CV is equal to ~0.995 for the first part, and ~0.25  for the second one. It is in good agreement with my result at LB:</p>\n\n<p>LB = ~(1/3*0.995 + 2/3*0.25) = ~0.5</p>\n\n<p>If it is as you say, i.e. practically all instances are well classified, you should obtain CV for no data leak part equal at least (1+0.5+0.333+0.25+0.2)/5 =~0.456 (it is the mean value for completely random sequence of 5 clusters in well classified instances) and your result at LB should exceed 0.636. </p>",
      "rawMarkdown": "[quote=MightyBird;119533]\r\nI found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction!\r\n[/quote]\r\n\r\n@MightyBird, I can't believe. I have divided the data set in two parts: one with data leak and second one containing the rest. My local CV is equal to ~0.995 for the first part, and ~0.25  for the second one. It is in good agreement with my result at LB:\r\n\r\nLB = ~(1/3*0.995 + 2/3*0.25) = ~0.5\r\n\r\nIf it is as you say, i.e. practically all instances are well classified, you should obtain CV for no data leak part equal at least (1+0.5+0.333+0.25+0.2)/5 =~0.456 (it is the mean value for completely random sequence of 5 clusters in well classified instances) and your result at LB should exceed 0.636.",
      "votes": null
    },
    {
      "id": "119679",
      "postDate": "05/12/2016 07:35:05",
      "content": "<p>Hi,</p>\n\n<p>I think that you cannot assume that the position in top 5 is evenly distributed. Maybe it's worse than that. Also note that I only use one sample of 1% of train.</p>\n\n<p>I think that you can only do one thing for sure. Get a lower limit for MAP@5. Which would be 0.1975 in this case.</p>\n\n<p>My best local score by the way was 0.4666 so far. And this includes the data leak information!</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nI think that you cannot assume that the position in top 5 is evenly distributed. Maybe it's worse than that. Also note that I only use one sample of 1% of train.\r\n\r\nI think that you can only do one thing for sure. Get a lower limit for MAP@5. Which would be 0.1975 in this case.\r\n\r\nMy best local score by the way was 0.4666 so far. And this includes the data leak information!\r\n\r\nGerhard",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119578,
      "author_name": "fsimond",
      "author_url": "",
      "post_date": "05/11/2016 14:07:40",
      "content": "<p>Thanks for the insights!</p>\n\n<p>I have a similar pipeline, except that I'm not getting such a good score, yet :)</p>\n\n<p>Did you only keep record with is_booking = 1 when creating your validation set ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119580,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/11/2016 14:38:15",
      "content": "<p>Yes, I selected 30.000 instances from the 2 year period where is_booking == 1. I did this because someone if not the Admin hinted that there is seasonality involved.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119676,
      "author_name": "sionek",
      "author_url": "",
      "post_date": "05/12/2016 07:09:38",
      "content": "<p>[quote=MightyBird;119533]\nI found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction!\n[/quote]</p>\n\n<p>@MightyBird, I can't believe. I have divided the data set in two parts: one with data leak and second one containing the rest. My local CV is equal to ~0.995 for the first part, and ~0.25  for the second one. It is in good agreement with my result at LB:</p>\n\n<p>LB = ~(1/3*0.995 + 2/3*0.25) = ~0.5</p>\n\n<p>If it is as you say, i.e. practically all instances are well classified, you should obtain CV for no data leak part equal at least (1+0.5+0.333+0.25+0.2)/5 =~0.456 (it is the mean value for completely random sequence of 5 clusters in well classified instances) and your result at LB should exceed 0.636. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119679,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/12/2016 07:35:05",
      "content": "<p>Hi,</p>\n\n<p>I think that you cannot assume that the position in top 5 is evenly distributed. Maybe it's worse than that. Also note that I only use one sample of 1% of train.</p>\n\n<p>I think that you can only do one thing for sure. Get a lower limit for MAP@5. Which would be 0.1975 in this case.</p>\n\n<p>My best local score by the way was 0.4666 so far. And this includes the data leak information!</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119533": "I learned so much from the Kaggle community that I wanted to share some information for this competition. Not too specific and not a script at the moment of course, but still.\r\n\r\nI did not yet reach 0.5 but only 0.00022 is missing...\r\n\r\nFist of all I created a train and a validation partition randomly (no time split!) using only 10% of the train data. I thought of creating further such sets to do cross validation but so far the single set served me well.\r\n\r\nThis needs around 3 minutes for a validation and evaluation run in my perl script per try. This allows for a lot of trial and error when looking for the right parameters.\r\n\r\nThe script takes 40 minutes at the moment for the complete data set. I run it usually 2-3 times per day, depending on the results of the validation runs.\r\n\r\nIn the script I use some of the data leak solutions but I also created some features/cluster sets on my own. This results in several sets of probable hotel clusters from which I can choose. Choosing and sorting of course is the crucial part.\r\n\r\nWhen doing a validation run I look at the results and how they were achieved. I evaluate hits and misses.\r\n\r\nI found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction! When looking at the misses I could mostly not see how I could have correctly forcasted them. So I think the rest of the work is either create new sets of possible clusters or rearrange the current predictions somehow.\r\n\r\nBy the way: I created another perl script that can merge submission files in a weighted way. So if you have two or more good submission files you possibly get a better one afterwards.\r\n\r\nCheers\r\n\r\nGerhard",
    "119578": "Thanks for the insights!\r\n\r\nI have a similar pipeline, except that I'm not getting such a good score, yet :)\r\n\r\nDid you only keep record with is_booking = 1 when creating your validation set ?",
    "119580": "Yes, I selected 30.000 instances from the 2 year period where is_booking == 1. I did this because someone if not the Admin hinted that there is seasonality involved.",
    "119676": "[quote=MightyBird;119533]\r\nI found that from my 30,000 validation instances only ~370 were completely misclassified. I was surprised. For all others the real hotel cluster was somewhere in the prediction!\r\n[/quote]\r\n\r\n@MightyBird, I can't believe. I have divided the data set in two parts: one with data leak and second one containing the rest. My local CV is equal to ~0.995 for the first part, and ~0.25  for the second one. It is in good agreement with my result at LB:\r\n\r\nLB = ~(1/3*0.995 + 2/3*0.25) = ~0.5\r\n\r\nIf it is as you say, i.e. practically all instances are well classified, you should obtain CV for no data leak part equal at least (1+0.5+0.333+0.25+0.2)/5 =~0.456 (it is the mean value for completely random sequence of 5 clusters in well classified instances) and your result at LB should exceed 0.636.",
    "119679": "Hi,\r\n\r\nI think that you cannot assume that the position in top 5 is evenly distributed. Maybe it's worse than that. Also note that I only use one sample of 1% of train.\r\n\r\nI think that you can only do one thing for sure. Get a lower limit for MAP@5. Which would be 0.1975 in this case.\r\n\r\nMy best local score by the way was 0.4666 so far. And this includes the data leak information!\r\n\r\nGerhard"
  },
  "source": "meta"
}