{
  "id": 21436,
  "title": "Because of the leakage? ",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21436",
  "author_name": "",
  "post_date": "2016-06-04T21:21:28.927Z",
  "votes": null,
  "comment_count": 3,
  "views": 708,
  "content": "<pre><code>[18]    eval-MAP@5:0.378607 train-MAP@5:0.402878\n[19]    eval-MAP@5:0.385755 train-MAP@5:0.410490\n[20]    eval-MAP@5:0.390674 train-MAP@5:0.415597\n[21]    eval-MAP@5:0.394828 train-MAP@5:0.420295\n[22]    eval-MAP@5:0.400880 train-MAP@5:0.426712\n[23]    eval-MAP@5:0.406655 train-MAP@5:0.432926\n[24]    eval-MAP@5:0.411265 train-MAP@5:0.437881\n</code></pre>\n\n<p>I use the xgboost to train the whole dataset with the evaluation metric (@CPMP, @dune_dweller posts on  <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric</a> ). The first 66% data as  training and the rest 34% as validation set. The validation score seems pretty high. Is it caused by information leak or there's something wrong with my code?</p>",
  "messages": [
    {
      "id": "122531",
      "postDate": "06/04/2016 21:21:28",
      "content": "<pre><code>[18]    eval-MAP@5:0.378607 train-MAP@5:0.402878\n[19]    eval-MAP@5:0.385755 train-MAP@5:0.410490\n[20]    eval-MAP@5:0.390674 train-MAP@5:0.415597\n[21]    eval-MAP@5:0.394828 train-MAP@5:0.420295\n[22]    eval-MAP@5:0.400880 train-MAP@5:0.426712\n[23]    eval-MAP@5:0.406655 train-MAP@5:0.432926\n[24]    eval-MAP@5:0.411265 train-MAP@5:0.437881\n</code></pre>\n\n<p>I use the xgboost to train the whole dataset with the evaluation metric (@CPMP, @dune_dweller posts on  <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric</a> ). The first 66% data as  training and the rest 34% as validation set. The validation score seems pretty high. Is it caused by information leak or there's something wrong with my code?</p>",
      "rawMarkdown": "[18]\teval-MAP@5:0.378607\ttrain-MAP@5:0.402878\r\n    [19]\teval-MAP@5:0.385755\ttrain-MAP@5:0.410490\r\n    [20]\teval-MAP@5:0.390674\ttrain-MAP@5:0.415597\r\n    [21]\teval-MAP@5:0.394828\ttrain-MAP@5:0.420295\r\n    [22]\teval-MAP@5:0.400880\ttrain-MAP@5:0.426712\r\n    [23]\teval-MAP@5:0.406655\ttrain-MAP@5:0.432926\r\n    [24]\teval-MAP@5:0.411265\ttrain-MAP@5:0.437881\r\n\r\nI use the xgboost to train the whole dataset with the evaluation metric (@CPMP, @dune_dweller posts on  https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric ). The first 66% data as  training and the rest 34% as validation set. The validation score seems pretty high. Is it caused by information leak or there's something wrong with my code?",
      "votes": null
    },
    {
      "id": "122543",
      "postDate": "06/05/2016 00:45:56",
      "content": "<p>That looks impressive. What is the score on LB?</p>",
      "rawMarkdown": "That looks impressive. What is the score on LB?",
      "votes": null
    },
    {
      "id": "122548",
      "postDate": "06/05/2016 02:18:44",
      "content": "<p>No. It extremely slow. I trained the whole data with is_booking==0 and 100+features .(66% data as train and 34% data as validation set). </p>",
      "rawMarkdown": "No. It extremely slow. I trained the whole data with is_booking==0 and 100+features .(66% data as train and 34% data as validation set).",
      "votes": null
    },
    {
      "id": "122571",
      "postDate": "06/05/2016 08:28:14",
      "content": "<p>My experience is that having user_id in easily leads to overfit.</p>",
      "rawMarkdown": "My experience is that having user_id in easily leads to overfit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122543,
      "author_name": "mandalsubhajit",
      "author_url": "",
      "post_date": "06/05/2016 00:45:56",
      "content": "<p>That looks impressive. What is the score on LB?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122548,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/05/2016 02:18:44",
      "content": "<p>No. It extremely slow. I trained the whole data with is_booking==0 and 100+features .(66% data as train and 34% data as validation set). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122571,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2016 08:28:14",
      "content": "<p>My experience is that having user_id in easily leads to overfit.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122531": "[18]\teval-MAP@5:0.378607\ttrain-MAP@5:0.402878\r\n    [19]\teval-MAP@5:0.385755\ttrain-MAP@5:0.410490\r\n    [20]\teval-MAP@5:0.390674\ttrain-MAP@5:0.415597\r\n    [21]\teval-MAP@5:0.394828\ttrain-MAP@5:0.420295\r\n    [22]\teval-MAP@5:0.400880\ttrain-MAP@5:0.426712\r\n    [23]\teval-MAP@5:0.406655\ttrain-MAP@5:0.432926\r\n    [24]\teval-MAP@5:0.411265\ttrain-MAP@5:0.437881\r\n\r\nI use the xgboost to train the whole dataset with the evaluation metric (@CPMP, @dune_dweller posts on  https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20556/map5-function-or-eval-metric ). The first 66% data as  training and the rest 34% as validation set. The validation score seems pretty high. Is it caused by information leak or there's something wrong with my code?",
    "122543": "That looks impressive. What is the score on LB?",
    "122548": "No. It extremely slow. I trained the whole data with is_booking==0 and 100+features .(66% data as train and 34% data as validation set).",
    "122571": "My experience is that having user_id in easily leads to overfit."
  },
  "source": "meta"
}