{
  "id": 21413,
  "title": "Team up",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21413",
  "author_name": "",
  "post_date": "2016-06-03T21:45:36.633Z",
  "votes": null,
  "comment_count": 11,
  "views": 1198,
  "content": "<p>I've worked only in ML solution (with xgboost) of the non leakage part of the problem.\nTraining with 2013 (date_time) year and hold 2014 (only the non leakage data relative to 2013) my CV score is 0.31 better than benchmark solutions (in 0.299 range).</p>\n\n<p>Interested in team up with complementary approach.</p>",
  "messages": [
    {
      "id": "122415",
      "postDate": "06/03/2016 21:45:36",
      "content": "<p>I've worked only in ML solution (with xgboost) of the non leakage part of the problem.\nTraining with 2013 (date_time) year and hold 2014 (only the non leakage data relative to 2013) my CV score is 0.31 better than benchmark solutions (in 0.299 range).</p>\n\n<p>Interested in team up with complementary approach.</p>",
      "rawMarkdown": "I've worked only in ML solution (with xgboost) of the non leakage part of the problem.\r\nTraining with 2013 (date_time) year and hold 2014 (only the non leakage data relative to 2013) my CV score is 0.31 better than benchmark solutions (in 0.299 range).\r\n\r\nInterested in team up with complementary approach.",
      "votes": null
    },
    {
      "id": "122417",
      "postDate": "06/03/2016 22:06:53",
      "content": "<p>I am interested ! I send you an email. </p>",
      "rawMarkdown": "I am interested ! I send you an email.",
      "votes": null
    },
    {
      "id": "122433",
      "postDate": "06/04/2016 01:46:06",
      "content": "<p>Does that mean that  there's also leakage data in train dataset? (2014 has some leakage for 2013?)</p>",
      "rawMarkdown": "Does that mean that  there's also leakage data in train dataset? (2014 has some leakage for 2013?)",
      "votes": null
    },
    {
      "id": "122472",
      "postDate": "06/04/2016 11:11:44",
      "content": "<p>Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).</p>",
      "rawMarkdown": "Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\r\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).",
      "votes": null
    },
    {
      "id": "122482",
      "postDate": "06/04/2016 14:45:32",
      "content": "<p>[quote=Jos&#233; A. Guerrero;122472]</p>\n\n<p>Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).</p>\n\n<p>[/quote]</p>\n\n<p>do you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?</p>",
      "rawMarkdown": "[quote=José A. Guerrero;122472]\r\n\r\nObviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\r\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).\r\n\r\n[/quote]\r\n\r\ndo you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?",
      "votes": null
    },
    {
      "id": "122487",
      "postDate": "06/04/2016 15:12:34",
      "content": "<p>@Jos&#233; A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.</p>",
      "rawMarkdown": "José A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.",
      "votes": null
    },
    {
      "id": "122495",
      "postDate": "06/04/2016 16:00:57",
      "content": "<p>[quote=Shubin;122482]</p>\n\n<p>[quote=Jos&#233; A. Guerrero;122472]</p>\n\n<p>Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).</p>\n\n<p>[/quote]</p>\n\n<p>do you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?</p>\n\n<p>[/quote]</p>\n\n<p>Not Shubin, only time split. Second half of 2014 (date_time) as holdout data set.</p>",
      "rawMarkdown": "[quote=Shubin;122482]\r\n\r\n[quote=José A. Guerrero;122472]\r\n\r\nObviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\r\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).\r\n\r\n[/quote]\r\n\r\ndo you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?\r\n\r\n\r\n[/quote]\r\n\r\nNot Shubin, only time split. Second half of 2014 (date_time) as holdout data set.",
      "votes": null
    },
    {
      "id": "122496",
      "postDate": "06/04/2016 16:01:18",
      "content": "<p>[quote=FengLi;122487]</p>\n\n<p>@Jos&#233; A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.</p>\n\n<p>[/quote]</p>\n\n<p>Yes.</p>",
      "rawMarkdown": "[quote=FengLi;122487]\r\n\r\n@José A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.\r\n\r\n[/quote]\r\n\r\nYes.",
      "votes": null
    },
    {
      "id": "122497",
      "postDate": "06/04/2016 16:03:34",
      "content": "<p>@Jos&#233; A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance</a></p>",
      "rawMarkdown": "José A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance",
      "votes": null
    },
    {
      "id": "122499",
      "postDate": "06/04/2016 16:14:26",
      "content": "<p>[quote=FengLi;122497]</p>\n\n<p>@Jos&#233; A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance</a></p>\n\n<p>[/quote]</p>\n\n<p>No. I'm using R.</p>",
      "rawMarkdown": "[quote=FengLi;122497]\r\n\r\n@José A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance\r\n\r\n[/quote]\r\n\r\nNo. I'm using R.",
      "votes": null
    },
    {
      "id": "122501",
      "postDate": "06/04/2016 16:24:24",
      "content": "<p>How long it takes you to train the whole data. For me it extremely slow with is_booking==1.</p>",
      "rawMarkdown": "How long it takes you to train the whole data. For me it extremely slow with is_booking==1.",
      "votes": null
    },
    {
      "id": "122502",
      "postDate": "06/04/2016 16:36:28",
      "content": "<p>Training this dataset with xgboost is slow. The solution: more time or more machine (or both!).\nI think the key is use clicks for feature engineering and then train only in bookings events.</p>",
      "rawMarkdown": "Training this dataset with xgboost is slow. The solution: more time or more machine (or both!).\r\nI think the key is use clicks for feature engineering and then train only in bookings events.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122417,
      "author_name": "alexandrearaujo",
      "author_url": "",
      "post_date": "06/03/2016 22:06:53",
      "content": "<p>I am interested ! I send you an email. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122433,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/04/2016 01:46:06",
      "content": "<p>Does that mean that  there's also leakage data in train dataset? (2014 has some leakage for 2013?)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122472,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2016 11:11:44",
      "content": "<p>Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122482,
      "author_name": "lishubin",
      "author_url": "",
      "post_date": "06/04/2016 14:45:32",
      "content": "<p>[quote=Jos&#233; A. Guerrero;122472]</p>\n\n<p>Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).</p>\n\n<p>[/quote]</p>\n\n<p>do you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122487,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/04/2016 15:12:34",
      "content": "<p>@Jos&#233; A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122495,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2016 16:00:57",
      "content": "<p>[quote=Shubin;122482]</p>\n\n<p>[quote=Jos&#233; A. Guerrero;122472]</p>\n\n<p>Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).</p>\n\n<p>[/quote]</p>\n\n<p>do you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?</p>\n\n<p>[/quote]</p>\n\n<p>Not Shubin, only time split. Second half of 2014 (date_time) as holdout data set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122496,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2016 16:01:18",
      "content": "<p>[quote=FengLi;122487]</p>\n\n<p>@Jos&#233; A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.</p>\n\n<p>[/quote]</p>\n\n<p>Yes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122497,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/04/2016 16:03:34",
      "content": "<p>@Jos&#233; A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122499,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2016 16:14:26",
      "content": "<p>[quote=FengLi;122497]</p>\n\n<p>@Jos&#233; A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance</a></p>\n\n<p>[/quote]</p>\n\n<p>No. I'm using R.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122501,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/04/2016 16:24:24",
      "content": "<p>How long it takes you to train the whole data. For me it extremely slow with is_booking==1.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122502,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "06/04/2016 16:36:28",
      "content": "<p>Training this dataset with xgboost is slow. The solution: more time or more machine (or both!).\nI think the key is use clicks for feature engineering and then train only in bookings events.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122415": "I've worked only in ML solution (with xgboost) of the non leakage part of the problem.\r\nTraining with 2013 (date_time) year and hold 2014 (only the non leakage data relative to 2013) my CV score is 0.31 better than benchmark solutions (in 0.299 range).\r\n\r\nInterested in team up with complementary approach.",
    "122417": "I am interested ! I send you an email.",
    "122433": "Does that mean that  there's also leakage data in train dataset? (2014 has some leakage for 2013?)",
    "122472": "Obviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\r\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).",
    "122482": "[quote=José A. Guerrero;122472]\r\n\r\nObviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\r\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).\r\n\r\n[/quote]\r\n\r\ndo you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?",
    "122487": "José A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.",
    "122495": "[quote=Shubin;122482]\r\n\r\n[quote=José A. Guerrero;122472]\r\n\r\nObviously. If you split the train by time in train and hold. The (user_location_city, orig_destination_distance) pairs in hold and seen in train are the leakage part. You can estimate the cv score in both (leakage and non leakage subsets of hold).\r\nJust now I'm fitting using last six month of 2014 as hold and the score in my non leakage subset is in 0.324 zone (Not submitted yet).\r\n\r\n[/quote]\r\n\r\ndo you take into account that you might have new user in your holdout dataset while in the actual leaderboard, all users have been seen.  is the cv the same?\r\n\r\n\r\n[/quote]\r\n\r\nNot Shubin, only time split. Second half of 2014 (date_time) as holdout data set.",
    "122496": "[quote=FengLi;122487]\r\n\r\n@José A. Guerrero So you just leave the leakage part and use ML learning method to the non-leakage part? Thank you.\r\n\r\n[/quote]\r\n\r\nYes.",
    "122497": "José A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance",
    "122499": "[quote=FengLi;122497]\r\n\r\n@José A. Guerrero Thank you. Do you know the reason that Jupyter just prints 37 validation information but the kernel and terminal is still working. I posted in https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/21428/question-about-printing-validation-performance\r\n\r\n[/quote]\r\n\r\nNo. I'm using R.",
    "122501": "How long it takes you to train the whole data. For me it extremely slow with is_booking==1.",
    "122502": "Training this dataset with xgboost is slow. The solution: more time or more machine (or both!).\r\nI think the key is use clicks for feature engineering and then train only in bookings events."
  },
  "source": "meta"
}