{
  "id": 306007,
  "title": "[Resolved] Question regarding evaluation metric",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/306007",
  "author_name": "Debarshi Chanda",
  "post_date": "2022-02-07T21:37:58.375000",
  "votes": 34,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I am not sure if I understand the metric correctly so I wanted to check my understanding<br>\nSuppose I have <code>ground_truth = [a, b, c, d, e]</code> and <code>preds = [b, c, a, d, f]</code> and suppose I want to calculate <code>AP@4</code></p>\n<pre><code>P@1 = 0/1,   rel@1 = 0\nP@2 = 1/2,   rel@2 = 0\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n</code></pre>\n<p>AP@4 = 1/4 x [(3/3) * 1 + (4/4) * 1] = 1/2</p>\n<p></p><hr><br>\n<strong>UPDATE</strong><br>\nCORRECT<p></p>\n<pre><code>P@1 = 1/1,   rel@1 = 1\nP@2 = 2/2,   rel@2 = 1\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n</code></pre>\n<p>AP@4 = 1/4 x [(1/1) * 1 + (2/2) * 1 + (3/3) * 1 + (4/4) * 1] = 1</p>",
  "messages": [
    {
      "id": 1680511,
      "postDate": "2022-02-07T21:37:58.377Z",
      "content": "<p>I am not sure if I understand the metric correctly so I wanted to check my understanding<br>\nSuppose I have <code>ground_truth = [a, b, c, d, e]</code> and <code>preds = [b, c, a, d, f]</code> and suppose I want to calculate <code>AP@4</code></p>\n<pre><code>P@1 = 0/1,   rel@1 = 0\nP@2 = 1/2,   rel@2 = 0\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n</code></pre>\n<p>AP@4 = 1/4 x [(3/3) * 1 + (4/4) * 1] = 1/2</p>\n<p></p><hr><br>\n<strong>UPDATE</strong><br>\nCORRECT<p></p>\n<pre><code>P@1 = 1/1,   rel@1 = 1\nP@2 = 2/2,   rel@2 = 1\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n</code></pre>\n<p>AP@4 = 1/4 x [(1/1) * 1 + (2/2) * 1 + (3/3) * 1 + (4/4) * 1] = 1</p>",
      "rawMarkdown": "I am not sure if I understand the metric correctly so I wanted to check my understanding\nSuppose I have `ground_truth = [a, b, c, d, e]` and `preds = [b, c, a, d, f]` and suppose I want to calculate `AP@4`\n```\nP@1 = 0/1,   rel@1 = 0\nP@2 = 1/2,   rel@2 = 0\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n```\n\nAP@4 = 1/4 x [(3/3) * 1 + (4/4) * 1] = 1/2\n\n<hr>\n**UPDATE**\nCORRECT\n```\nP@1 = 1/1,   rel@1 = 1\nP@2 = 2/2,   rel@2 = 1\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n```\n\nAP@4 = 1/4 x [(1/1) * 1 + (2/2) * 1 + (3/3) * 1 + (4/4) * 1] = 1",
      "votes": 33
    },
    {
      "id": 1683460,
      "postDate": "2022-02-09T20:11:55.890Z",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Should the formula be revised at the evaluation page to this<br>\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$<br>\ninstead of this<br>\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$</p>",
      "rawMarkdown": "@inversion Should the formula be revised at the evaluation page to this\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$\ninstead of this\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$",
      "votes": 13,
      "replies": [
        {
          "id": 1684317,
          "postDate": "2022-02-10T12:33:55.150Z",
          "content": "<p>I agree with <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> <br>\nbecause of this line: <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py#L39\" target=\"_blank\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py#L39</a></p>",
          "rawMarkdown": "I agree with @debarshichanda \nbecause of this line: https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py#L39\n",
          "votes": 4
        },
        {
          "id": 1693194,
          "postDate": "2022-02-16T14:03:32.790Z",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> Can anyone confirm this?</p>",
          "rawMarkdown": "@inversion @maggiemd Can anyone confirm this?",
          "votes": 1
        },
        {
          "id": 1722781,
          "postDate": "2022-03-14T21:41:33.110Z",
          "content": "<p>Sorry for missing this! The metric page has been updated. Please see this thread.</p>\n<p><a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194</a></p>",
          "rawMarkdown": "Sorry for missing this! The metric page has been updated. Please see this thread.\n\nhttps://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194"
        },
        {
          "id": 1722804,
          "postDate": "2022-03-14T21:54:54.517Z",
          "content": "<p>Clicking the link says - \"No Access\"</p>",
          "rawMarkdown": "Clicking the link says - \"No Access\""
        },
        {
          "id": 1722826,
          "postDate": "2022-03-14T22:27:34.230Z",
          "content": "<p><a href=\"https://www.kaggle.com/atulverma\" target=\"_blank\">@atulverma</a> Should work now. :-)</p>",
          "rawMarkdown": "@atulverma Should work now. :-)"
        },
        {
          "id": 1722879,
          "postDate": "2022-03-15T00:55:29.020Z",
          "content": "<p>Sorry but if <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> below (<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573</a>) is right, then formula on evaluation page doesn't mirror that </p>",
          "rawMarkdown": "Sorry but if @cdeotte below (https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573) is right, then formula on evaluation page doesn't mirror that "
        },
        {
          "id": 1722890,
          "postDate": "2022-03-15T01:03:07.670Z",
          "content": "<blockquote>\n  <p>Sorry but if <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> below (<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573</a>) is right, then formula doesn't mirror that</p>\n</blockquote>\n<p>Which formula is incorrect. The formula that Debarshi proposes is different than mine and the competition's metric. Debarshi uses <code>min(n,12)</code> twice when it should be <code>min(m,12)</code> and <code>min(n,12)</code>. Where one refers to number of predictions and one refers to number of ground truths. (Also note that <code>12</code> is the <code>k</code> in <code>@k</code>).</p>",
          "rawMarkdown": ">Sorry but if @cdeotte below (https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573) is right, then formula doesn't mirror that\n\nWhich formula is incorrect. The formula that Debarshi proposes is different than mine and the competition's metric. Debarshi uses `min(n,12)` twice when it should be `min(m,12)` and `min(n,12)`. Where one refers to number of predictions and one refers to number of ground truths. (Also note that `12` is the `k` in `@k`)."
        },
        {
          "id": 1722942,
          "postDate": "2022-03-15T01:44:25.467Z",
          "content": "<p>Both.. but i was referring to your formula here <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194</a> which has been reflected in the evaluation page now.</p>\n<p>I think</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(m, 12)}  \\sum_{k=1}^{min(n,12)} \\frac{1}{min(k, n)} rel(k)$$</p>\n<p>Or</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(m, 12)}  \\sum_{k=1}^{min(n,12)} P(k)$$ </p>\n<p>Where</p>\n<p>$$P(k) = \\frac{1}{min(k, n)} rel(k)$$</p>",
          "rawMarkdown": "Both.. but i was referring to your formula here https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194 which has been reflected in the evaluation page now.\n\n\nI think\n\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(m, 12)}  \\sum_{k=1}^{min(n,12)} \\frac{1}{min(k, n)} rel(k)$$\n\nOr\n\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(m, 12)}  \\sum_{k=1}^{min(n,12)} P(k)$$ \n\nWhere\n\n$$P(k) = \\frac{1}{min(k, n)} rel(k)$$"
        },
        {
          "id": 1722957,
          "postDate": "2022-03-15T02:08:52.780Z",
          "content": "<p>Or</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} (\\sum_{k=1}^{min(n,12)} \\frac{1}{k} rel(k)) \\times rel(k) $$</p>\n<p>Or</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$</p>\n<p>Where</p>\n<p>$$P(k) = \\sum_{k=1}^{min(n,12)} rel(k) \\frac{1}{k} $$</p>",
          "rawMarkdown": "Or\n\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} (\\sum_{k=1}^{min(n,12)} \\frac{1}{k} rel(k)) \\times rel(k) $$\n\nOr\n\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$\n\nWhere\n\n$$P(k) = \\sum_{k=1}^{min(n,12)} rel(k) \\frac{1}{k} $$"
        },
        {
          "id": 1722959,
          "postDate": "2022-03-15T02:12:34.737Z",
          "content": "<p>Why do you think it is wrong? Can you provide an example?</p>\n<p>(Note that your two formulas don't include precision, they only include <code>rel(k)</code>. Since the metric is <code>mAP</code> i.e. mean average precision, we need to include the computation of precision)</p>",
          "rawMarkdown": "Why do you think it is wrong? Can you provide an example?\n\n(Note that your two formulas don't include precision, they only include `rel(k)`. Since the metric is `mAP` i.e. mean average precision, we need to include the computation of precision)"
        },
        {
          "id": 1723012,
          "postDate": "2022-03-15T03:44:16.207Z",
          "content": "<p>Below you say this</p>\n<blockquote>\n  <p>Precision@1 = 0/1, recall doesn't increase<br>\n  Precision@2 = 1/2, recall increase<br>\n  Precision@3 = 2/3, recall increase<br>\n  Precision@4 = 3/4, recall increase</p>\n</blockquote>\n<p>numerator is sum of rel(k)..  one summation is missing in the current formula<br>\ndenominator is min(k,n) where k is 1 to min(n,12) .. this division is also missing in the current formula.</p>\n<p>Both the formulas in this post <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1722957\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1722957</a> include Precision as defined in your post below</p>",
          "rawMarkdown": "Below you say this\n\n> Precision@1 = 0/1, recall doesn't increase\n> Precision@2 = 1/2, recall increase\n> Precision@3 = 2/3, recall increase\n> Precision@4 = 3/4, recall increase\n\nnumerator is sum of rel(k)..  one summation is missing in the current formula\ndenominator is min(k,n) where k is 1 to min(n,12) .. this division is also missing in the current formula.\n\nBoth the formulas in this post https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1722957 include Precision as defined in your post below"
        },
        {
          "id": 1723024,
          "postDate": "2022-03-15T04:01:26.970Z",
          "content": "<p>The new kaggle formula is now correct. In my example, we have</p>\n<p><code>(0/1 * 0) + (1/2 * 1) + (2/3 * 1) + (3/4 * 1)</code> where this is <code>(p(0) * rel(0)) + (p(1) * rel(1)) + etc etc</code></p>\n<p>This is the summation of <code>p(k) * rel(k) for k in [1, 2, ..., n]</code> where <code>n</code> is the number of predictions. After computing this summation, we divide by <code>m</code> where <code>m</code> is the number of ground truths. (And we use 12 in place of <code>n</code> and <code>m</code> if it is lower than either).</p>\n<p>In your formula, you need to multiply <code>p(k)</code> times <code>rel(k)</code>. The function <code>rel(k)</code> is a mask that only considers certain <code>p(k)</code></p>",
          "rawMarkdown": "The new kaggle formula is now correct. In my example, we have\n\n`(0/1 * 0) + (1/2 * 1) + (2/3 * 1) + (3/4 * 1)` where this is `(p(0) * rel(0)) + (p(1) * rel(1)) + etc etc`\n\nThis is the summation of `p(k) * rel(k) for k in [1, 2, ..., n]` where `n` is the number of predictions. After computing this summation, we divide by `m` where `m` is the number of ground truths. (And we use 12 in place of `n` and `m` if it is lower than either).\n\nIn your formula, you need to multiply `p(k)` times `rel(k)`. The function `rel(k)` is a mask that only considers certain `p(k)`",
          "votes": 1,
          "replies": [
            {
              "id": 1723027,
              "postDate": "2022-03-15T04:06:37.023Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 1723036,
              "postDate": "2022-03-15T04:22:42.023Z",
              "content": "<blockquote>\n  <p>In your formula, you need to multiply <code>p(k)</code> times <code>rel(k)</code>. The function <code>rel(k)</code> is a mask that only considers certain <code>p(k)</code></p>\n</blockquote>\n<p>Yes, I missed that.. but the current formula doesn't describe P(k) and P(k), it seems also uses rel(k) which is what I am trying above (Updated)</p>\n<p>It seems P(k) is summation of rel(k) divided by k.</p>",
              "rawMarkdown": "> In your formula, you need to multiply `p(k)` times `rel(k)`. The function `rel(k)` is a mask that only considers certain `p(k)`\n\nYes, I missed that.. but the current formula doesn't describe P(k) and P(k), it seems also uses rel(k) which is what I am trying above (Updated)\n\nIt seems P(k) is summation of rel(k) divided by k."
            }
          ]
        },
        {
          "id": 1723041,
          "postDate": "2022-03-15T04:33:51.977Z",
          "content": "<p><code>p(k)</code> is precision. It is <code>TP/ (TP+FP)</code> when only considering your first <code>k</code> predictions. It is completely different than <code>rel(k)</code>. The term <code>rel(k)</code> is 1 if you newest prediction (i.e. prediction at place <code>k</code>)  improves your recall. And 0 if your newest prediction does not improve your recall.</p>\n<p>I agree it is very weird. It makes more sense if you look at precision recall plots (where <code>rel(k)</code> moves the plot line to the right whereas p(k) only moves plot line up and down)</p>",
          "rawMarkdown": "`p(k)` is precision. It is `TP/ (TP+FP)` when only considering your first `k` predictions. It is completely different than `rel(k)`. The term `rel(k)` is 1 if you newest prediction (i.e. prediction at place `k`)  improves your recall. And 0 if your newest prediction does not improve your recall.\n\nI agree it is very weird. It makes more sense if you look at precision recall plots (where `rel(k)` moves the plot line to the right whereas p(k) only moves plot line up and down)"
        }
      ]
    },
    {
      "id": 1680513,
      "postDate": "2022-02-07T21:42:43.057Z",
      "content": "<p>You can also walk through this code:</p>\n<p><a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py</a></p>",
      "rawMarkdown": "You can also walk through this code:\n\nhttps://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py",
      "votes": 8,
      "replies": [
        {
          "id": 1680546,
          "postDate": "2022-02-07T22:15:26.457Z",
          "content": "<p>Thanks for this but I am not sure why the order doesn't matter<br>\nIf <code>ground_truth = [a, b, c, d, e]</code> then the above code gives the same answer for both <code>[d, c, b, a, f]</code> and <code>[a, b, c, d, e]</code> i.e. 1 for AP@4</p>",
          "rawMarkdown": "Thanks for this but I am not sure why the order doesn't matter\nIf `ground_truth = [a, b, c, d, e]` then the above code gives the same answer for both `[d, c, b, a, f]` and `[a, b, c, d, e]` i.e. 1 for AP@4",
          "votes": 2
        },
        {
          "id": 1680595,
          "postDate": "2022-02-07T23:22:36.720Z",
          "content": "<p>There is no order to the ground truth, it only matters whether your predictions are found in the ground truth. For both of your predictions, the first 4 labels are found in the ground truth. Since you are looking at AP@4, they both score <code>1.0</code>.</p>",
          "rawMarkdown": "There is no order to the ground truth, it only matters whether your predictions are found in the ground truth. For both of your predictions, the first 4 labels are found in the ground truth. Since you are looking at AP@4, they both score `1.0`.",
          "votes": 6
        },
        {
          "id": 1680604,
          "postDate": "2022-02-07T23:41:58.867Z",
          "content": "<p>Thank you so much! Understood my mistake</p>",
          "rawMarkdown": "Thank you so much! Understood my mistake",
          "votes": 1
        },
        {
          "id": 1687848,
          "postDate": "2022-02-13T07:49:20.147Z",
          "content": "<blockquote>\n  <p>You can also walk through this code:<br>\n  <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py</a></p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> that code won't handle well a case when <code>y_true=[]</code> i.e. a costumer doesn't buy anything on the test-data timeframe.</p>\n<p>Example:</p>\n<pre><code>{'Customer01': {\n    'y_true': [], \n    'y_pred': [1,2,3],\n}},\n{'Customer02': {\n    'y_true': [1,2,3], \n    'y_pred': [1,2,3],\n}},\n\nObtained Score = 0.5\nExpected Score = 1.0\n</code></pre>\n<p>Citing from Evaluation section in the competition overview:</p>\n<blockquote>\n  <ul>\n  <li>Customer that did not make any purchase during test period are excluded from the scoring.</li>\n  <li>There is never a penalty for using the full 12 predictions for a customer that ordered fewer than 12 items; thus, it's advantageous to make 12 predictions for each customer.</li>\n  </ul>\n</blockquote>\n<p>Cheers,</p>",
          "rawMarkdown": "> You can also walk through this code:\nhttps://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\n\n@inversion that code won't handle well a case when `y_true=[]` i.e. a costumer doesn't buy anything on the test-data timeframe.\n\nExample:\n```\n{'Customer01': {\n    'y_true': [], \n    'y_pred': [1,2,3],\n}},\n{'Customer02': {\n    'y_true': [1,2,3], \n    'y_pred': [1,2,3],\n}},\n\nObtained Score = 0.5\nExpected Score = 1.0\n```\n\nCiting from Evaluation section in the competition overview:\n> - Customer that did not make any purchase during test period are excluded from the scoring.\n- There is never a penalty for using the full 12 predictions for a customer that ordered fewer than 12 items; thus, it's advantageous to make 12 predictions for each customer.\n\nCheers,"
        },
        {
          "id": 1688869,
          "postDate": "2022-02-13T21:33:32.150Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1693397,
          "postDate": "2022-02-16T16:43:41.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> <br>\nhow can AP@4 would be same for the two predictions as ordering effect average precision. </p>",
          "rawMarkdown": "@debarshichanda \nhow can AP@4 would be same for the two predictions as ordering effect average precision. "
        },
        {
          "id": 1693573,
          "postDate": "2022-02-16T18:52:21.570Z",
          "content": "<p>The order of your predictions matters but order of ground truth does not matter. In your original question <code>ground_truth = [a, b, c, d, e] and preds = [b, c, a, d, f]</code>, the answer is <code>AP@4 = 1</code>. </p>\n<p>(Note that the competition metric is <code>mAP@12</code> not <code>mAP@4</code> but we can do <code>mAP@4</code> in example below, we must just stop after using the first 4 predictions)</p>\n<p>If your predictions are <code>[f, b, c, a, d]</code>, then your <code>AP@4</code> is less than 1. Note that we don't use the last prediction of <code>d</code> here since we are doing <code>AP@4</code>.</p>\n<p>Precision@1 = 0/1, recall doesn't increase<br>\nPrecision@2 = 1/2, recall increase<br>\nPrecision@3 = 2/3, recall increase<br>\nPrecision@4 = 3/4, recall increase</p>\n<p><code>AP@4 = sum of \"recall increase\" / min(4,len(true)) = (1/2 + 2/3 + 3/4)/4 = 23/48 = 0.479</code></p>",
          "rawMarkdown": "The order of your predictions matters but order of ground truth does not matter. In your original question `ground_truth = [a, b, c, d, e] and preds = [b, c, a, d, f]`, the answer is `AP@4 = 1`. \n\n(Note that the competition metric is `mAP@12` not `mAP@4` but we can do `mAP@4` in example below, we must just stop after using the first 4 predictions)\n\nIf your predictions are `[f, b, c, a, d]`, then your `AP@4` is less than 1. Note that we don't use the last prediction of `d` here since we are doing `AP@4`.\n\nPrecision@1 = 0/1, recall doesn't increase\nPrecision@2 = 1/2, recall increase\nPrecision@3 = 2/3, recall increase\nPrecision@4 = 3/4, recall increase\n\n`AP@4 = sum of \"recall increase\" / min(4,len(true)) = (1/2 + 2/3 + 3/4)/4 = 23/48 = 0.479`",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1683460,
      "author_name": "Debarshi Chanda",
      "author_url": "",
      "post_date": "2022-02-09T20:11:55.890000",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Should the formula be revised at the evaluation page to this<br>\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$<br>\ninstead of this<br>\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1684317,
          "author_name": "chuan",
          "author_url": "",
          "post_date": "2022-02-10T12:33:55.150000",
          "content": "<p>I agree with <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> <br>\nbecause of this line: <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py#L39\" target=\"_blank\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py#L39</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1693194,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2022-02-16T14:03:32.790000",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> Can anyone confirm this?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1722781,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2022-03-14T21:41:33.110000",
          "content": "<p>Sorry for missing this! The metric page has been updated. Please see this thread.</p>\n<p><a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722804,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-14T21:54:54.517000",
          "content": "<p>Clicking the link says - \"No Access\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722826,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2022-03-14T22:27:34.230000",
          "content": "<p><a href=\"https://www.kaggle.com/atulverma\" target=\"_blank\">@atulverma</a> Should work now. :-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722879,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-15T00:55:29.020000",
          "content": "<p>Sorry but if <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> below (<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573</a>) is right, then formula on evaluation page doesn't mirror that </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722890,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-15T01:03:07.670000",
          "content": "<blockquote>\n  <p>Sorry but if <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> below (<a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1693573</a>) is right, then formula doesn't mirror that</p>\n</blockquote>\n<p>Which formula is incorrect. The formula that Debarshi proposes is different than mine and the competition's metric. Debarshi uses <code>min(n,12)</code> twice when it should be <code>min(m,12)</code> and <code>min(n,12)</code>. Where one refers to number of predictions and one refers to number of ground truths. (Also note that <code>12</code> is the <code>k</code> in <code>@k</code>).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722942,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-15T01:44:25.467000",
          "content": "<p>Both.. but i was referring to your formula here <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/309152#1711194</a> which has been reflected in the evaluation page now.</p>\n<p>I think</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(m, 12)}  \\sum_{k=1}^{min(n,12)} \\frac{1}{min(k, n)} rel(k)$$</p>\n<p>Or</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(m, 12)}  \\sum_{k=1}^{min(n,12)} P(k)$$ </p>\n<p>Where</p>\n<p>$$P(k) = \\frac{1}{min(k, n)} rel(k)$$</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722957,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-15T02:08:52.780000",
          "content": "<p>Or</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} (\\sum_{k=1}^{min(n,12)} \\frac{1}{k} rel(k)) \\times rel(k) $$</p>\n<p>Or</p>\n<p>$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k) $$</p>\n<p>Where</p>\n<p>$$P(k) = \\sum_{k=1}^{min(n,12)} rel(k) \\frac{1}{k} $$</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722959,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-15T02:12:34.737000",
          "content": "<p>Why do you think it is wrong? Can you provide an example?</p>\n<p>(Note that your two formulas don't include precision, they only include <code>rel(k)</code>. Since the metric is <code>mAP</code> i.e. mean average precision, we need to include the computation of precision)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1723012,
          "author_name": "AtulVerma",
          "author_url": "",
          "post_date": "2022-03-15T03:44:16.207000",
          "content": "<p>Below you say this</p>\n<blockquote>\n  <p>Precision@1 = 0/1, recall doesn't increase<br>\n  Precision@2 = 1/2, recall increase<br>\n  Precision@3 = 2/3, recall increase<br>\n  Precision@4 = 3/4, recall increase</p>\n</blockquote>\n<p>numerator is sum of rel(k)..  one summation is missing in the current formula<br>\ndenominator is min(k,n) where k is 1 to min(n,12) .. this division is also missing in the current formula.</p>\n<p>Both the formulas in this post <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1722957\" target=\"_blank\">https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306007#1722957</a> include Precision as defined in your post below</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1723024,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-15T04:01:26.970000",
          "content": "<p>The new kaggle formula is now correct. In my example, we have</p>\n<p><code>(0/1 * 0) + (1/2 * 1) + (2/3 * 1) + (3/4 * 1)</code> where this is <code>(p(0) * rel(0)) + (p(1) * rel(1)) + etc etc</code></p>\n<p>This is the summation of <code>p(k) * rel(k) for k in [1, 2, ..., n]</code> where <code>n</code> is the number of predictions. After computing this summation, we divide by <code>m</code> where <code>m</code> is the number of ground truths. (And we use 12 in place of <code>n</code> and <code>m</code> if it is lower than either).</p>\n<p>In your formula, you need to multiply <code>p(k)</code> times <code>rel(k)</code>. The function <code>rel(k)</code> is a mask that only considers certain <code>p(k)</code></p>",
          "votes": 1,
          "replies": [
            {
              "id": 1723027,
              "author_name": "",
              "author_url": "",
              "post_date": "2022-03-15T04:06:37.023000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1723036,
              "author_name": "AtulVerma",
              "author_url": "",
              "post_date": "2022-03-15T04:22:42.023000",
              "content": "<blockquote>\n  <p>In your formula, you need to multiply <code>p(k)</code> times <code>rel(k)</code>. The function <code>rel(k)</code> is a mask that only considers certain <code>p(k)</code></p>\n</blockquote>\n<p>Yes, I missed that.. but the current formula doesn't describe P(k) and P(k), it seems also uses rel(k) which is what I am trying above (Updated)</p>\n<p>It seems P(k) is summation of rel(k) divided by k.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1723041,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-03-15T04:33:51.977000",
          "content": "<p><code>p(k)</code> is precision. It is <code>TP/ (TP+FP)</code> when only considering your first <code>k</code> predictions. It is completely different than <code>rel(k)</code>. The term <code>rel(k)</code> is 1 if you newest prediction (i.e. prediction at place <code>k</code>)  improves your recall. And 0 if your newest prediction does not improve your recall.</p>\n<p>I agree it is very weird. It makes more sense if you look at precision recall plots (where <code>rel(k)</code> moves the plot line to the right whereas p(k) only moves plot line up and down)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1680513,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2022-02-07T21:42:43.057000",
      "content": "<p>You can also walk through this code:</p>\n<p><a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py</a></p>",
      "votes": 8,
      "replies": [
        {
          "id": 1680546,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2022-02-07T22:15:26.457000",
          "content": "<p>Thanks for this but I am not sure why the order doesn't matter<br>\nIf <code>ground_truth = [a, b, c, d, e]</code> then the above code gives the same answer for both <code>[d, c, b, a, f]</code> and <code>[a, b, c, d, e]</code> i.e. 1 for AP@4</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1680595,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2022-02-07T23:22:36.720000",
          "content": "<p>There is no order to the ground truth, it only matters whether your predictions are found in the ground truth. For both of your predictions, the first 4 labels are found in the ground truth. Since you are looking at AP@4, they both score <code>1.0</code>.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1680604,
          "author_name": "Debarshi Chanda",
          "author_url": "",
          "post_date": "2022-02-07T23:41:58.867000",
          "content": "<p>Thank you so much! Understood my mistake</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1687848,
          "author_name": "Thariq Nugrohotomo",
          "author_url": "",
          "post_date": "2022-02-13T07:49:20.147000",
          "content": "<blockquote>\n  <p>You can also walk through this code:<br>\n  <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py\" target=\"_blank\">https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py</a></p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> that code won't handle well a case when <code>y_true=[]</code> i.e. a costumer doesn't buy anything on the test-data timeframe.</p>\n<p>Example:</p>\n<pre><code>{'Customer01': {\n    'y_true': [], \n    'y_pred': [1,2,3],\n}},\n{'Customer02': {\n    'y_true': [1,2,3], \n    'y_pred': [1,2,3],\n}},\n\nObtained Score = 0.5\nExpected Score = 1.0\n</code></pre>\n<p>Citing from Evaluation section in the competition overview:</p>\n<blockquote>\n  <ul>\n  <li>Customer that did not make any purchase during test period are excluded from the scoring.</li>\n  <li>There is never a penalty for using the full 12 predictions for a customer that ordered fewer than 12 items; thus, it's advantageous to make 12 predictions for each customer.</li>\n  </ul>\n</blockquote>\n<p>Cheers,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1688869,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-13T21:33:32.150000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1693397,
          "author_name": "ayush",
          "author_url": "",
          "post_date": "2022-02-16T16:43:41.740000",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> <br>\nhow can AP@4 would be same for the two predictions as ordering effect average precision. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1693573,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-02-16T18:52:21.570000",
          "content": "<p>The order of your predictions matters but order of ground truth does not matter. In your original question <code>ground_truth = [a, b, c, d, e] and preds = [b, c, a, d, f]</code>, the answer is <code>AP@4 = 1</code>. </p>\n<p>(Note that the competition metric is <code>mAP@12</code> not <code>mAP@4</code> but we can do <code>mAP@4</code> in example below, we must just stop after using the first 4 predictions)</p>\n<p>If your predictions are <code>[f, b, c, a, d]</code>, then your <code>AP@4</code> is less than 1. Note that we don't use the last prediction of <code>d</code> here since we are doing <code>AP@4</code>.</p>\n<p>Precision@1 = 0/1, recall doesn't increase<br>\nPrecision@2 = 1/2, recall increase<br>\nPrecision@3 = 2/3, recall increase<br>\nPrecision@4 = 3/4, recall increase</p>\n<p><code>AP@4 = sum of \"recall increase\" / min(4,len(true)) = (1/2 + 2/3 + 3/4)/4 = 23/48 = 0.479</code></p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1680511": "I am not sure if I understand the metric correctly so I wanted to check my understanding\nSuppose I have `ground_truth = [a, b, c, d, e]` and `preds = [b, c, a, d, f]` and suppose I want to calculate `AP@4`\n```\nP@1 = 0/1,   rel@1 = 0\nP@2 = 1/2,   rel@2 = 0\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n```\n\nAP@4 = 1/4 x [(3/3) * 1 + (4/4) * 1] = 1/2\n\n<hr>\n**UPDATE**\nCORRECT\n```\nP@1 = 1/1,   rel@1 = 1\nP@2 = 2/2,   rel@2 = 1\nP@3 = 3/3,   rel@3 = 1\nP@4 = 4/4,   rel@4 = 1\n```\n\nAP@4 = 1/4 x [(1/1) * 1 + (2/2) * 1 + (3/3) * 1 + (4/4) * 1] = 1",
    "1683460": "@inversion Should the formula be revised at the evaluation page to this\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\frac{1}{min(n, 12)}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$\ninstead of this\n$$MAP@12 = \\frac{1}{U} \\sum_{u=1}^{U}  \\sum_{k=1}^{min(n,12)} P(k) \\times rel(k)$$",
    "1680513": "You can also walk through this code:\n\nhttps://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/average_precision.py"
  }
}