{
  "id": 20182,
  "title": "Does MAP@5 averaged by # of predictions?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20182",
  "author_name": "",
  "post_date": "2016-04-16T20:21:51.587Z",
  "votes": 1,
  "comment_count": 6,
  "views": 1530,
  "content": "<p>I have a question about the evaluation rule:</p>\n\n<p>MAP@5 in the evaluation rule is the sum of P(k), which is not divided by the # of prediction, aka sum of P(K)/min(n,5). Since MAP@5 usually is divided by 5, I need to confirm the formula in evaluation rule is not a typo.</p>\n\n<p>If that's not a typo, submitting 5 predictions for each event will always be the better strategy than submitting less predictions, right? Can someone confirm this?</p>",
  "messages": [
    {
      "id": "115172",
      "postDate": "04/16/2016 20:21:51",
      "content": "<p>I have a question about the evaluation rule:</p>\n\n<p>MAP@5 in the evaluation rule is the sum of P(k), which is not divided by the # of prediction, aka sum of P(K)/min(n,5). Since MAP@5 usually is divided by 5, I need to confirm the formula in evaluation rule is not a typo.</p>\n\n<p>If that's not a typo, submitting 5 predictions for each event will always be the better strategy than submitting less predictions, right? Can someone confirm this?</p>",
      "rawMarkdown": "I have a question about the evaluation rule:\r\n\r\nMAP@5 in the evaluation rule is the sum of P(k), which is not divided by the # of prediction, aka sum of P(K)/min(n,5). Since MAP@5 usually is divided by 5, I need to confirm the formula in evaluation rule is not a typo.\r\n\r\nIf that's not a typo, submitting 5 predictions for each event will always be the better strategy than submitting less predictions, right? Can someone confirm this?",
      "votes": null
    },
    {
      "id": "115181",
      "postDate": "04/16/2016 22:54:16",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "115453",
      "postDate": "04/18/2016 18:42:28",
      "content": "<p>@Chao Pan, </p>\n\n<p>Good catch - but it's not a typo. The term of 1/min(m,5) is still there, but it's eliminated because m is always 1.  Since m is the actual # of hotel clusters that is booked by the user. </p>\n\n<p>Btw, our production code for MAP can be found <a href=\"https://www.kaggle.com/wiki/MeanAveragePrecision\">here</a>. Another <a href=\"https://www.kaggle.com/c/coupon-purchase-prediction/details/evaluation\">example</a> of using MAP@K but with m &gt;= 1. </p>\n\n<p>EDIT: changed n to m to avoid confusion</p>",
      "rawMarkdown": "Chao Pan, \r\n\r\nGood catch - but it's not a typo. The term of 1/min(m,5) is still there, but it's eliminated because m is always 1.  Since m is the actual # of hotel clusters that is booked by the user. \r\n\r\nBtw, our production code for MAP can be found [here][1]. Another [example][2] of using MAP@K but with m >= 1. \r\n\r\nEDIT: changed n to m to avoid confusion\r\n\r\n\r\n  [1]: https://www.kaggle.com/wiki/MeanAveragePrecision\r\n  [2]: https://www.kaggle.com/c/coupon-purchase-prediction/details/evaluation",
      "votes": null
    },
    {
      "id": "115557",
      "postDate": "04/19/2016 05:35:12",
      "content": "<p>@Wendy Kan,</p>\n\n<p>It does matter order?\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\nI know this metric was used in coupon prediction, but I completely forgot it.\nPlease look at attached the png.\nThanks.</p>",
      "rawMarkdown": "Wendy Kan,\r\n\r\nIt does matter order?\r\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\r\nI know this metric was used in coupon prediction, but I completely forgot it.\r\nPlease look at attached the png.\r\nThanks.",
      "votes": null
    },
    {
      "id": "115599",
      "postDate": "04/19/2016 09:25:40",
      "content": "<p>[quote=ONODERA;115557]</p>\n\n<p>@Wendy Kan,</p>\n\n<p>It does matter order?\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\nI know this metric was used in coupon prediction, but I completely forgot it.\nPlease look at attached the png.\nThanks.</p>\n\n<p>[/quote]</p>\n\n<p>I think this is how it works..\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.</p>\n\n<p>For a particular booking.. if the answer is.. for example:  17..</p>\n\n<p>If you submit:</p>\n\n<p>a) 17,3,55,23,54:  AP = 1/1 = 1\nb) 43,56,17,12,4: AP = 1/3 = 0.333\nc) 43,1,46,13,17: AP = 1/5 = 0.2</p>\n\n<p>So here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..</p>\n\n<p>Your average precision for each booking is: 1/position of actual cluster in your submission.\nThe average precision is just 0 if you don't get the actual correct answer. </p>\n\n<p>So you should always submit five answers, but order them based on likelihood!</p>",
      "rawMarkdown": "[quote=ONODERA;115557]\r\n\r\n@Wendy Kan,\r\n\r\nIt does matter order?\r\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\r\nI know this metric was used in coupon prediction, but I completely forgot it.\r\nPlease look at attached the png.\r\nThanks.\r\n\r\n[/quote]\r\n\r\nI think this is how it works..\r\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.\r\n\r\nFor a particular booking.. if the answer is.. for example:  17..\r\n\r\nIf you submit:\r\n\r\na) 17,3,55,23,54:  AP = 1/1 = 1\r\nb) 43,56,17,12,4: AP = 1/3 = 0.333\r\nc) 43,1,46,13,17: AP = 1/5 = 0.2\r\n\r\nSo here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..\r\n\r\nYour average precision for each booking is: 1/position of actual cluster in your submission.\r\nThe average precision is just 0 if you don't get the actual correct answer. \r\n\r\nSo you should always submit five answers, but order them based on likelihood!",
      "votes": null
    },
    {
      "id": "115606",
      "postDate": "04/19/2016 09:48:05",
      "content": "<p>[quote=Dylan;115599]</p>\n\n<p>I think this is how it works..\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.</p>\n\n<p>For a particular booking.. if the answer is.. for example:  17..</p>\n\n<p>If you submit:</p>\n\n<p>a) 17,3,55,23,54:  AP = 1/1 = 1\nb) 43,56,17,12,4: AP = 1/3 = 0.333\nc) 43,1,46,13,17: AP = 1/5 = 0.2</p>\n\n<p>So here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..</p>\n\n<p>Your average precision for each booking is: 1/position of actual cluster in your submission.\nThe average precision is just 0 if you don't get the actual correct answer. </p>\n\n<p>So you should always submit five answers, but order them based on likelihood!</p>\n\n<p>[/quote]</p>\n\n<p>My understanding is below,</p>\n\n<p>Order does matter, but in this competition we don't need to worry about it.</p>\n\n<p>Because answer(k) is always 1.</p>\n\n<p>Anyway, it seems we should submit 5 answers to each id by ordering them based on probability.</p>\n\n<p>Thank you. I think I get it.</p>",
      "rawMarkdown": "[quote=Dylan;115599]\r\n\r\n\r\nI think this is how it works..\r\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.\r\n\r\nFor a particular booking.. if the answer is.. for example:  17..\r\n\r\nIf you submit:\r\n\r\na) 17,3,55,23,54:  AP = 1/1 = 1\r\nb) 43,56,17,12,4: AP = 1/3 = 0.333\r\nc) 43,1,46,13,17: AP = 1/5 = 0.2\r\n\r\nSo here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..\r\n\r\nYour average precision for each booking is: 1/position of actual cluster in your submission.\r\nThe average precision is just 0 if you don't get the actual correct answer. \r\n\r\nSo you should always submit five answers, but order them based on likelihood!\r\n\r\n[/quote]\r\n\r\nMy understanding is below,\r\n\r\nOrder does matter, but in this competition we don't need to worry about it.\r\n\r\nBecause answer(k) is always 1.\r\n\r\nAnyway, it seems we should submit 5 answers to each id by ordering them based on probability.\r\n\r\nThank you. I think I get it.",
      "votes": null
    },
    {
      "id": "115696",
      "postDate": "04/19/2016 16:56:17",
      "content": "<p>@Dylan is correct. Order does matter, so put your most likely hotel prediction first! </p>\n\n<p>I thought the best way to show it is to write a script. So <a href=\"https://www.kaggle.com/wendykan/expedia-hotel-recommendations/map-k-demo/notebook\">here</a> it is. </p>",
      "rawMarkdown": "Dylan is correct. Order does matter, so put your most likely hotel prediction first! \r\n\r\nI thought the best way to show it is to write a script. So [here][1] it is. \r\n\r\n\r\n  [1]: https://www.kaggle.com/wendykan/expedia-hotel-recommendations/map-k-demo/notebook",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 115181,
      "author_name": "jasontam",
      "author_url": "",
      "post_date": "04/16/2016 22:54:16",
      "content": "<p>yes</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115453,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/18/2016 18:42:28",
      "content": "<p>@Chao Pan, </p>\n\n<p>Good catch - but it's not a typo. The term of 1/min(m,5) is still there, but it's eliminated because m is always 1.  Since m is the actual # of hotel clusters that is booked by the user. </p>\n\n<p>Btw, our production code for MAP can be found <a href=\"https://www.kaggle.com/wiki/MeanAveragePrecision\">here</a>. Another <a href=\"https://www.kaggle.com/c/coupon-purchase-prediction/details/evaluation\">example</a> of using MAP@K but with m &gt;= 1. </p>\n\n<p>EDIT: changed n to m to avoid confusion</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115557,
      "author_name": "onodera",
      "author_url": "",
      "post_date": "04/19/2016 05:35:12",
      "content": "<p>@Wendy Kan,</p>\n\n<p>It does matter order?\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\nI know this metric was used in coupon prediction, but I completely forgot it.\nPlease look at attached the png.\nThanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115599,
      "author_name": "kellybus",
      "author_url": "",
      "post_date": "04/19/2016 09:25:40",
      "content": "<p>[quote=ONODERA;115557]</p>\n\n<p>@Wendy Kan,</p>\n\n<p>It does matter order?\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\nI know this metric was used in coupon prediction, but I completely forgot it.\nPlease look at attached the png.\nThanks.</p>\n\n<p>[/quote]</p>\n\n<p>I think this is how it works..\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.</p>\n\n<p>For a particular booking.. if the answer is.. for example:  17..</p>\n\n<p>If you submit:</p>\n\n<p>a) 17,3,55,23,54:  AP = 1/1 = 1\nb) 43,56,17,12,4: AP = 1/3 = 0.333\nc) 43,1,46,13,17: AP = 1/5 = 0.2</p>\n\n<p>So here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..</p>\n\n<p>Your average precision for each booking is: 1/position of actual cluster in your submission.\nThe average precision is just 0 if you don't get the actual correct answer. </p>\n\n<p>So you should always submit five answers, but order them based on likelihood!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115606,
      "author_name": "onodera",
      "author_url": "",
      "post_date": "04/19/2016 09:48:05",
      "content": "<p>[quote=Dylan;115599]</p>\n\n<p>I think this is how it works..\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.</p>\n\n<p>For a particular booking.. if the answer is.. for example:  17..</p>\n\n<p>If you submit:</p>\n\n<p>a) 17,3,55,23,54:  AP = 1/1 = 1\nb) 43,56,17,12,4: AP = 1/3 = 0.333\nc) 43,1,46,13,17: AP = 1/5 = 0.2</p>\n\n<p>So here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..</p>\n\n<p>Your average precision for each booking is: 1/position of actual cluster in your submission.\nThe average precision is just 0 if you don't get the actual correct answer. </p>\n\n<p>So you should always submit five answers, but order them based on likelihood!</p>\n\n<p>[/quote]</p>\n\n<p>My understanding is below,</p>\n\n<p>Order does matter, but in this competition we don't need to worry about it.</p>\n\n<p>Because answer(k) is always 1.</p>\n\n<p>Anyway, it seems we should submit 5 answers to each id by ordering them based on probability.</p>\n\n<p>Thank you. I think I get it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115696,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "04/19/2016 16:56:17",
      "content": "<p>@Dylan is correct. Order does matter, so put your most likely hotel prediction first! </p>\n\n<p>I thought the best way to show it is to write a script. So <a href=\"https://www.kaggle.com/wendykan/expedia-hotel-recommendations/map-k-demo/notebook\">here</a> it is. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "115172": "I have a question about the evaluation rule:\r\n\r\nMAP@5 in the evaluation rule is the sum of P(k), which is not divided by the # of prediction, aka sum of P(K)/min(n,5). Since MAP@5 usually is divided by 5, I need to confirm the formula in evaluation rule is not a typo.\r\n\r\nIf that's not a typo, submitting 5 predictions for each event will always be the better strategy than submitting less predictions, right? Can someone confirm this?",
    "115181": "yes",
    "115453": "Chao Pan, \r\n\r\nGood catch - but it's not a typo. The term of 1/min(m,5) is still there, but it's eliminated because m is always 1.  Since m is the actual # of hotel clusters that is booked by the user. \r\n\r\nBtw, our production code for MAP can be found [here][1]. Another [example][2] of using MAP@K but with m >= 1. \r\n\r\nEDIT: changed n to m to avoid confusion\r\n\r\n\r\n  [1]: https://www.kaggle.com/wiki/MeanAveragePrecision\r\n  [2]: https://www.kaggle.com/c/coupon-purchase-prediction/details/evaluation",
    "115557": "Wendy Kan,\r\n\r\nIt does matter order?\r\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\r\nI know this metric was used in coupon prediction, but I completely forgot it.\r\nPlease look at attached the png.\r\nThanks.",
    "115599": "[quote=ONODERA;115557]\r\n\r\n@Wendy Kan,\r\n\r\nIt does matter order?\r\nSuppose actual is [a,b,c,d,e], and if prediction is [e,d,c,b,a], the score is 0.2?\r\nI know this metric was used in coupon prediction, but I completely forgot it.\r\nPlease look at attached the png.\r\nThanks.\r\n\r\n[/quote]\r\n\r\nI think this is how it works..\r\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.\r\n\r\nFor a particular booking.. if the answer is.. for example:  17..\r\n\r\nIf you submit:\r\n\r\na) 17,3,55,23,54:  AP = 1/1 = 1\r\nb) 43,56,17,12,4: AP = 1/3 = 0.333\r\nc) 43,1,46,13,17: AP = 1/5 = 0.2\r\n\r\nSo here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..\r\n\r\nYour average precision for each booking is: 1/position of actual cluster in your submission.\r\nThe average precision is just 0 if you don't get the actual correct answer. \r\n\r\nSo you should always submit five answers, but order them based on likelihood!",
    "115606": "[quote=Dylan;115599]\r\n\r\n\r\nI think this is how it works..\r\nOrder does matter. There is only one correct answer (correct hotel cluster) per booking.\r\n\r\nFor a particular booking.. if the answer is.. for example:  17..\r\n\r\nIf you submit:\r\n\r\na) 17,3,55,23,54:  AP = 1/1 = 1\r\nb) 43,56,17,12,4: AP = 1/3 = 0.333\r\nc) 43,1,46,13,17: AP = 1/5 = 0.2\r\n\r\nSo here, submitting (a) gives you precision of 1, submitting (c) only gives 0.2..\r\n\r\nYour average precision for each booking is: 1/position of actual cluster in your submission.\r\nThe average precision is just 0 if you don't get the actual correct answer. \r\n\r\nSo you should always submit five answers, but order them based on likelihood!\r\n\r\n[/quote]\r\n\r\nMy understanding is below,\r\n\r\nOrder does matter, but in this competition we don't need to worry about it.\r\n\r\nBecause answer(k) is always 1.\r\n\r\nAnyway, it seems we should submit 5 answers to each id by ordering them based on probability.\r\n\r\nThank you. I think I get it.",
    "115696": "Dylan is correct. Order does matter, so put your most likely hotel prediction first! \r\n\r\nI thought the best way to show it is to write a script. So [here][1] it is. \r\n\r\n\r\n  [1]: https://www.kaggle.com/wendykan/expedia-hotel-recommendations/map-k-demo/notebook"
  },
  "source": "meta"
}