{
  "id": 364530,
  "title": "Which metric is correct?",
  "url": "/competitions/otto-recommender-system/discussion/364530",
  "author_name": "",
  "post_date": "2022-11-07T03:46:33.156204200Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The evaluation metric in this competition is Recall@k. In this case, which is correct, (1) or (2)?</p>\n<blockquote>\n  <p>label = {1: [111], 2:[111,222]}<br>\n  predict = {1: [999,888], 2:[111,222}</p>\n</blockquote>\n<p>(1) Average Recall@k per session<br>\n(0.0 + 1.0) / 2 = 0.5</p>\n<p>(2) Divide \"the total number of hits\" by \"the total number of events\"<br>\n2 / 3 = 0.66…</p>\n<p>In this <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">great notebook</a>  by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> CV score is calculated with (1). However, I think the <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">host github</a> has the metric of (2).</p>\n<p>Sorry if I misunderstood.<br>\n<a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>, I would be glad to know.</p>",
  "messages": [
    {
      "id": "2019867",
      "postDate": "11/07/2022 03:46:33",
      "content": "<p>The evaluation metric in this competition is Recall@k. In this case, which is correct, (1) or (2)?</p>\n<blockquote>\n  <p>label = {1: [111], 2:[111,222]}<br>\n  predict = {1: [999,888], 2:[111,222}</p>\n</blockquote>\n<p>(1) Average Recall@k per session<br>\n(0.0 + 1.0) / 2 = 0.5</p>\n<p>(2) Divide \"the total number of hits\" by \"the total number of events\"<br>\n2 / 3 = 0.66…</p>\n<p>In this <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">great notebook</a>  by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> CV score is calculated with (1). However, I think the <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">host github</a> has the metric of (2).</p>\n<p>Sorry if I misunderstood.<br>\n<a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>, I would be glad to know.</p>",
      "rawMarkdown": "The evaluation metric in this competition is Recall@k. In this case, which is correct, (1) or (2)?\n\n> label = {1: [111], 2:[111,222]}\npredict = {1: [999,888], 2:[111,222}\n\n(1) Average Recall@k per session\n(0.0 + 1.0) / 2 = 0.5\n\n(2) Divide \"the total number of hits\" by \"the total number of events\"\n2 / 3 = 0.66...\n\nIn this [great notebook](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework)  by @radek1 CV score is calculated with (1). However, I think the [host github](https://github.com/otto-de/recsys-dataset) has the metric of (2).\n\nSorry if I misunderstood.\n@pnormann, I would be glad to know.",
      "votes": null
    },
    {
      "id": "2019909",
      "postDate": "11/07/2022 04:21:17",
      "content": "<p>Yes, you are right, <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a>! I just spotted this today 🙂</p>\n<p>Please see the discussion in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534\" target=\"_blank\">this thread</a>.</p>\n<p>It turns out it's number 2!</p>\n<p>(though of course would be nice to hear confirmation that indeed this is how it is also implemented on kaggle 😉)</p>",
      "rawMarkdown": "Yes, you are right, @dehokanta! I just spotted this today 🙂\n\nPlease see the discussion in [this thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534).\n\nIt turns out it's number 2!\n\n(though of course would be nice to hear confirmation that indeed this is how it is also implemented on kaggle 😉)",
      "votes": null
    },
    {
      "id": "2020239",
      "postDate": "11/07/2022 09:49:19",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ,Thanks for your comment.😊</p>\n<p>Yes, it would be nice if the LB metric could be clarified. <br>\nLooking forward to <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> response!</p>",
      "rawMarkdown": "radek1 ,Thanks for your comment.😊\n\nYes, it would be nice if the LB metric could be clarified. \nLooking forward to @pnormann response!",
      "votes": null
    },
    {
      "id": "2020630",
      "postDate": "11/07/2022 16:45:14",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a>, thanks for your question! I briefly chatted with <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> and confirmed that the Kaggle LB implementation also uses number (2) for the metric calculation 😌</p>",
      "rawMarkdown": "Hey @radek1 and @dehokanta, thanks for your question! I briefly chatted with @inversion and confirmed that the Kaggle LB implementation also uses number (2) for the metric calculation 😌",
      "votes": null
    },
    {
      "id": "2020800",
      "postDate": "11/07/2022 19:33:16",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>! 🙂 Thanks a lot for following up on this and sharing this with us 🙏</p>",
      "rawMarkdown": "Thank you @pnormann! 🙂 Thanks a lot for following up on this and sharing this with us 🙏",
      "votes": null
    },
    {
      "id": "2020879",
      "postDate": "11/07/2022 20:40:49",
      "content": "<p>Thanks for the quick response. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> 😊</p>",
      "rawMarkdown": "Thanks for the quick response. @pnormann 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2019909,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/07/2022 04:21:17",
      "content": "<p>Yes, you are right, <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a>! I just spotted this today 🙂</p>\n<p>Please see the discussion in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534\" target=\"_blank\">this thread</a>.</p>\n<p>It turns out it's number 2!</p>\n<p>(though of course would be nice to hear confirmation that indeed this is how it is also implemented on kaggle 😉)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2020239,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "11/07/2022 09:49:19",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ,Thanks for your comment.😊</p>\n<p>Yes, it would be nice if the LB metric could be clarified. <br>\nLooking forward to <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> response!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2020630,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/07/2022 16:45:14",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/dehokanta\" target=\"_blank\">@dehokanta</a>, thanks for your question! I briefly chatted with <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> and confirmed that the Kaggle LB implementation also uses number (2) for the metric calculation 😌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2020800,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/07/2022 19:33:16",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>! 🙂 Thanks a lot for following up on this and sharing this with us 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2020879,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "11/07/2022 20:40:49",
          "content": "<p>Thanks for the quick response. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2019867": "The evaluation metric in this competition is Recall@k. In this case, which is correct, (1) or (2)?\n\n> label = {1: [111], 2:[111,222]}\npredict = {1: [999,888], 2:[111,222}\n\n(1) Average Recall@k per session\n(0.0 + 1.0) / 2 = 0.5\n\n(2) Divide \"the total number of hits\" by \"the total number of events\"\n2 / 3 = 0.66...\n\nIn this [great notebook](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework)  by @radek1 CV score is calculated with (1). However, I think the [host github](https://github.com/otto-de/recsys-dataset) has the metric of (2).\n\nSorry if I misunderstood.\n@pnormann, I would be glad to know.",
    "2019909": "Yes, you are right, @dehokanta! I just spotted this today 🙂\n\nPlease see the discussion in [this thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364534).\n\nIt turns out it's number 2!\n\n(though of course would be nice to hear confirmation that indeed this is how it is also implemented on kaggle 😉)",
    "2020239": "radek1 ,Thanks for your comment.😊\n\nYes, it would be nice if the LB metric could be clarified. \nLooking forward to @pnormann response!",
    "2020630": "Hey @radek1 and @dehokanta, thanks for your question! I briefly chatted with @inversion and confirmed that the Kaggle LB implementation also uses number (2) for the metric calculation 😌",
    "2020800": "Thank you @pnormann! 🙂 Thanks a lot for following up on this and sharing this with us 🙏",
    "2020879": "Thanks for the quick response. @pnormann 😊"
  },
  "source": "meta"
}