{
  "id": 21463,
  "title": "benchmarks without the leak? ",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21463",
  "author_name": "",
  "post_date": "2016-06-06T13:43:45.817Z",
  "votes": null,
  "comment_count": 6,
  "views": 1091,
  "content": "<p>Hi all... </p>\n\n<p>I was wondering if anyone had any benchmark scores that didn't take advantage of the leaks?  I understand from a competition perspective this is silly,  but I am trying to get a sense of how much the leak contributes to the overall mapk score on the leaderboard... or, more importantly, what's reasonable to expect without exploiting the leak. </p>\n\n<p>thanks. </p>",
  "messages": [
    {
      "id": "122682",
      "postDate": "06/06/2016 13:43:45",
      "content": "<p>Hi all... </p>\n\n<p>I was wondering if anyone had any benchmark scores that didn't take advantage of the leaks?  I understand from a competition perspective this is silly,  but I am trying to get a sense of how much the leak contributes to the overall mapk score on the leaderboard... or, more importantly, what's reasonable to expect without exploiting the leak. </p>\n\n<p>thanks. </p>",
      "rawMarkdown": "Hi all... \r\n\r\nI was wondering if anyone had any benchmark scores that didn't take advantage of the leaks?  I understand from a competition perspective this is silly,  but I am trying to get a sense of how much the leak contributes to the overall mapk score on the leaderboard... or, more importantly, what's reasonable to expect without exploiting the leak. \r\n\r\nthanks.",
      "votes": null
    },
    {
      "id": "122725",
      "postDate": "06/06/2016 20:19:14",
      "content": "<p>On my cv set (981939 id's), I get:</p>\n\n<ul>\n<li>0.910973 (for the 322065 leaked id's)</li>\n<li>0.3232433 (for the 659874 nonleaked id's)</li>\n</ul>\n\n<p>To obtain 0.516012 in the full validation set... (0.5070 in the LB)</p>",
      "rawMarkdown": "On my cv set (981939 id's), I get:\r\n\r\n* 0.910973 (for the 322065 leaked id's)\r\n* 0.3232433 (for the 659874 nonleaked id's)\r\n\r\nTo obtain 0.516012 in the full validation set... (0.5070 in the LB)",
      "votes": null
    },
    {
      "id": "122764",
      "postDate": "06/07/2016 01:20:39",
      "content": "<p>@Manel The leaked id is just based on org_destination_distance ? And when you trained the model, you include the leakage id or remove it?  Thank you.</p>",
      "rawMarkdown": "Manel The leaked id is just based on org_destination_distance ? And when you trained the model, you include the leakage id or remove it?  Thank you.",
      "votes": null
    },
    {
      "id": "122766",
      "postDate": "06/07/2016 01:56:07",
      "content": "<p>@Manel, how are you identifying the leaked id's?</p>",
      "rawMarkdown": "Manel, how are you identifying the leaked id's?",
      "votes": null
    },
    {
      "id": "122769",
      "postDate": "06/07/2016 02:22:21",
      "content": "<p>@manel -- thanks. that was enough information to convince me that I still had a bug in my code, and I was able to ferret it out. </p>",
      "rawMarkdown": "manel -- thanks. that was enough information to convince me that I still had a bug in my code, and I was able to ferret it out.",
      "votes": null
    },
    {
      "id": "122791",
      "postDate": "06/07/2016 06:41:45",
      "content": "<p>@FengLi You have to read the comments by @adam in the first permanent post. It is very well explained.\nYes, I consider all the training set to predict the unleaked data.</p>\n\n<p>@Igor Lins e Silva I just run the part considering only the data leak (orig_dest_distance,...) and then I see the id's that I have filled in the test, these are what I said leaked id's and the rest unleaked</p>",
      "rawMarkdown": "FengLi You have to read the comments by @adam in the first permanent post. It is very well explained.\r\nYes, I consider all the training set to predict the unleaked data.\r\n\r\n@Igor Lins e Silva I just run the part considering only the data leak (orig_dest_distance,...) and then I see the id's that I have filled in the test, these are what I said leaked id's and the rest unleaked",
      "votes": null
    },
    {
      "id": "122834",
      "postDate": "06/07/2016 16:43:51",
      "content": "<p>@Igor , you might also want to check out this script which find out the leak ratio\n<a href=\"https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/output\">https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/output</a></p>\n\n<p>[quote=Igor Lins e Silva;122766]</p>\n\n<p>@Manel, how are you identifying the leaked id's?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Igor , you might also want to check out this script which find out the leak ratio\r\nhttps://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/output\r\n\r\n[quote=Igor Lins e Silva;122766]\r\n\r\n@Manel, how are you identifying the leaked id's?\r\n\r\n[/quote]",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122725,
      "author_name": "manels",
      "author_url": "",
      "post_date": "06/06/2016 20:19:14",
      "content": "<p>On my cv set (981939 id's), I get:</p>\n\n<ul>\n<li>0.910973 (for the 322065 leaked id's)</li>\n<li>0.3232433 (for the 659874 nonleaked id's)</li>\n</ul>\n\n<p>To obtain 0.516012 in the full validation set... (0.5070 in the LB)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122764,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/07/2016 01:20:39",
      "content": "<p>@Manel The leaked id is just based on org_destination_distance ? And when you trained the model, you include the leakage id or remove it?  Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122766,
      "author_name": "igorls",
      "author_url": "",
      "post_date": "06/07/2016 01:56:07",
      "content": "<p>@Manel, how are you identifying the leaked id's?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122769,
      "author_name": "knitcode",
      "author_url": "",
      "post_date": "06/07/2016 02:22:21",
      "content": "<p>@manel -- thanks. that was enough information to convince me that I still had a bug in my code, and I was able to ferret it out. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122791,
      "author_name": "manels",
      "author_url": "",
      "post_date": "06/07/2016 06:41:45",
      "content": "<p>@FengLi You have to read the comments by @adam in the first permanent post. It is very well explained.\nYes, I consider all the training set to predict the unleaked data.</p>\n\n<p>@Igor Lins e Silva I just run the part considering only the data leak (orig_dest_distance,...) and then I see the id's that I have filled in the test, these are what I said leaked id's and the rest unleaked</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122834,
      "author_name": "paullo0106",
      "author_url": "",
      "post_date": "06/07/2016 16:43:51",
      "content": "<p>@Igor , you might also want to check out this script which find out the leak ratio\n<a href=\"https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/output\">https://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/output</a></p>\n\n<p>[quote=Igor Lins e Silva;122766]</p>\n\n<p>@Manel, how are you identifying the leaked id's?</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122682": "Hi all... \r\n\r\nI was wondering if anyone had any benchmark scores that didn't take advantage of the leaks?  I understand from a competition perspective this is silly,  but I am trying to get a sense of how much the leak contributes to the overall mapk score on the leaderboard... or, more importantly, what's reasonable to expect without exploiting the leak. \r\n\r\nthanks.",
    "122725": "On my cv set (981939 id's), I get:\r\n\r\n* 0.910973 (for the 322065 leaked id's)\r\n* 0.3232433 (for the 659874 nonleaked id's)\r\n\r\nTo obtain 0.516012 in the full validation set... (0.5070 in the LB)",
    "122764": "Manel The leaked id is just based on org_destination_distance ? And when you trained the model, you include the leakage id or remove it?  Thank you.",
    "122766": "Manel, how are you identifying the leaked id's?",
    "122769": "manel -- thanks. that was enough information to convince me that I still had a bug in my code, and I was able to ferret it out.",
    "122791": "FengLi You have to read the comments by @adam in the first permanent post. It is very well explained.\r\nYes, I consider all the training set to predict the unleaked data.\r\n\r\n@Igor Lins e Silva I just run the part considering only the data leak (orig_dest_distance,...) and then I see the id's that I have filled in the test, these are what I said leaked id's and the rest unleaked",
    "122834": "Igor , you might also want to check out this script which find out the leak ratio\r\nhttps://www.kaggle.com/sionek/expedia-hotel-recommendations/simple-validation/output\r\n\r\n[quote=Igor Lins e Silva;122766]\r\n\r\n@Manel, how are you identifying the leaked id's?\r\n\r\n[/quote]"
  },
  "source": "meta"
}