{
  "id": 307041,
  "title": "Competition metric (MAP) code",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307041",
  "author_name": "",
  "post_date": "2022-02-12T08:13:29.670525700Z",
  "votes": 32,
  "comment_count": 4,
  "views": 0,
  "content": "<pre><code>import numpy as np\n\n\ndef average_precision(data_true: dict, data_predicted: dict, k: int, norm: bool = True) -&gt; float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if not data_true:\n        raise ValueError('data_true is empty')\n\n    average_precision_sum = 0.0\n\n    for key, items_true in data_true.items():\n        items_predicted = data_predicted.get(key, [])\n\n        n_items_true = len(items_true)\n        n_items_predicted = min(len(items_predicted), k)\n\n        if n_items_true == 0 or n_items_predicted == 0:\n            continue\n\n        n_correct_items = 0\n        precision = 0.0\n\n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            average_precision_sum += precision / min(n_items_true, k)\n        else:\n            average_precision_sum += precision / k\n\n    return average_precision_sum / len(data_true)\n</code></pre>\n<p>This competition has Mean <strong>Normalized</strong> average precision score as a metric, so <code>norm</code> should be set to <code>True</code></p>",
  "messages": [
    {
      "id": "1686637",
      "postDate": "02/12/2022 08:13:29",
      "content": "<pre><code>import numpy as np\n\n\ndef average_precision(data_true: dict, data_predicted: dict, k: int, norm: bool = True) -&gt; float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if not data_true:\n        raise ValueError('data_true is empty')\n\n    average_precision_sum = 0.0\n\n    for key, items_true in data_true.items():\n        items_predicted = data_predicted.get(key, [])\n\n        n_items_true = len(items_true)\n        n_items_predicted = min(len(items_predicted), k)\n\n        if n_items_true == 0 or n_items_predicted == 0:\n            continue\n\n        n_correct_items = 0\n        precision = 0.0\n\n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            average_precision_sum += precision / min(n_items_true, k)\n        else:\n            average_precision_sum += precision / k\n\n    return average_precision_sum / len(data_true)\n</code></pre>\n<p>This competition has Mean <strong>Normalized</strong> average precision score as a metric, so <code>norm</code> should be set to <code>True</code></p>",
      "rawMarkdown": "```\nimport numpy as np\n\n\ndef average_precision(data_true: dict, data_predicted: dict, k: int, norm: bool = True) -> float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if not data_true:\n        raise ValueError('data_true is empty')\n\n    average_precision_sum = 0.0\n\n    for key, items_true in data_true.items():\n        items_predicted = data_predicted.get(key, [])\n\n        n_items_true = len(items_true)\n        n_items_predicted = min(len(items_predicted), k)\n\n        if n_items_true == 0 or n_items_predicted == 0:\n            continue\n\n        n_correct_items = 0\n        precision = 0.0\n\n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            average_precision_sum += precision / min(n_items_true, k)\n        else:\n            average_precision_sum += precision / k\n\n    return average_precision_sum / len(data_true)\n```\n\nThis competition has Mean **Normalized** average precision score as a metric, so `norm` should be set to `True`",
      "votes": null
    },
    {
      "id": "1689676",
      "postDate": "02/14/2022 11:55:38",
      "content": "<p>Thank you for sharing, <a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a>! Did you do tests 😄?</p>",
      "rawMarkdown": "Thank you for sharing, @nroman! Did you do tests 😄?",
      "votes": null
    },
    {
      "id": "1689909",
      "postDate": "02/14/2022 15:23:33",
      "content": "<p><a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a><br>\nFor the following code, output is 0.5. 🤔:</p>\n<pre><code>k = 3\nsample_data = {\"a\": [\"1\",\"2\",\"3\"], \"b\": [] }\n\naverage_precision(data_true=sample_data, data_predicted=sample_data, k)\n</code></pre>\n<p>Perhaps you need to replace this…</p>\n<pre><code>if n_items_true == 0 or n_items_predicted == 0:\n    continue\n</code></pre>\n<p>with this?</p>\n<pre><code>if n_items_true == 0:\n    del data_true[key]\n    continue\nif n_items_predicted == 0:\n    continue\n</code></pre>",
      "rawMarkdown": "nroman\nFor the following code, output is 0.5. 🤔:\n```\nk = 3\nsample_data = {\"a\": [\"1\",\"2\",\"3\"], \"b\": [] }\n\naverage_precision(data_true=sample_data, data_predicted=sample_data, k)\n```\n   \n    \n   \nPerhaps you need to replace this...\n```\nif n_items_true == 0 or n_items_predicted == 0:\n    continue\n```\nwith this?\n```\nif n_items_true == 0:\n    del data_true[key]\n    continue\nif n_items_predicted == 0:\n    continue\n```",
      "votes": null
    },
    {
      "id": "1689952",
      "postDate": "02/14/2022 15:52:09",
      "content": "<p>for those that prefer pandas/apply, here's an exact implementation of the <a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a>'s </p>\n<pre><code>def pandas_average_precision(\n    data_true: pd.Series,\n    data_predicted: pd.Series,\n    k: int,\n    norm: bool = True, \n) -&gt; float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if len(data_true)==0:\n        raise ValueError('data_true is empty')\n\n    # convert to df so we can use `apply`\n    eval_df = pd.DataFrame({\"true\": data_true, \"predicted\": data_predicted})\n\n    # replace na predictions with empty list\n    eval_df[\"predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: x if isinstance(x, list) else []\n    )\n\n    # getting the counts of true/predicted\n    eval_df[\"n_items_true\"] = eval_df[\"true\"].apply(len)\n    eval_df[\"n_items_predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: min(len(x), k)\n    )\n\n    # ignore zero true or predicted\n    non_zero_filter = (eval_df[\"n_items_true\"] &gt; 0) &amp; (eval_df[\"n_items_predicted\"] &gt; 0)\n    eval_df = eval_df[non_zero_filter].copy()\n\n    def row_precision(items_true, items_predicted, n_items_predicted, n_items_true, norm=norm):\n        n_correct_items = 0\n        precision = 0.0\n\n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            return precision / min(n_items_true, k)\n        else:\n            return precision / k\n\n    eval_df[\"row_precision\"] = eval_df.apply(\n        lambda x: row_precision(\n            x[\"true\"], x[\"predicted\"], x[\"n_items_predicted\"], x[\"n_items_true\"]\n        ), axis=1\n    )\n\n    return eval_df[\"row_precision\"].sum() / len(data_true)\n</code></pre>",
      "rawMarkdown": "for those that prefer pandas/apply, here's an exact implementation of the @nroman's \n\n```\ndef pandas_average_precision(\n    data_true: pd.Series,\n    data_predicted: pd.Series,\n    k: int,\n    norm: bool = True, \n) -> float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if len(data_true)==0:\n        raise ValueError('data_true is empty')\n\n    # convert to df so we can use `apply`\n    eval_df = pd.DataFrame({\"true\": data_true, \"predicted\": data_predicted})\n    \n    # replace na predictions with empty list\n    eval_df[\"predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: x if isinstance(x, list) else []\n    )\n    \n    # getting the counts of true/predicted\n    eval_df[\"n_items_true\"] = eval_df[\"true\"].apply(len)\n    eval_df[\"n_items_predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: min(len(x), k)\n    )\n    \n    # ignore zero true or predicted\n    non_zero_filter = (eval_df[\"n_items_true\"] > 0) & (eval_df[\"n_items_predicted\"] > 0)\n    eval_df = eval_df[non_zero_filter].copy()\n    \n    def row_precision(items_true, items_predicted, n_items_predicted, n_items_true, norm=norm):\n        n_correct_items = 0\n        precision = 0.0\n        \n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            return precision / min(n_items_true, k)\n        else:\n            return precision / k\n            \n    eval_df[\"row_precision\"] = eval_df.apply(\n        lambda x: row_precision(\n            x[\"true\"], x[\"predicted\"], x[\"n_items_predicted\"], x[\"n_items_true\"]\n        ), axis=1\n    )\n    \n    return eval_df[\"row_precision\"].sum() / len(data_true)\n```",
      "votes": null
    },
    {
      "id": "1698820",
      "postDate": "02/20/2022 17:35:28",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a>, I think your modification is right, because customers having no purchases during test period are excluded from scoring. Hence, they should be excluded here to remove the effect on denominator <code>len(data_true)</code>. And, I think this issue is similar to what I comment in <a href=\"https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12\" target=\"_blank\">this kernel</a>.</p>\n<p>If there's any misunderstanding, please correct me. Thanks a lot.</p>",
      "rawMarkdown": "Hi @jacob34, I think your modification is right, because customers having no purchases during test period are excluded from scoring. Hence, they should be excluded here to remove the effect on denominator `len(data_true)`. And, I think this issue is similar to what I comment in [this kernel](https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12).\n\nIf there's any misunderstanding, please correct me. Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1689676,
      "author_name": "vad13irt",
      "author_url": "",
      "post_date": "02/14/2022 11:55:38",
      "content": "<p>Thank you for sharing, <a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a>! Did you do tests 😄?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1689909,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "02/14/2022 15:23:33",
      "content": "<p><a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a><br>\nFor the following code, output is 0.5. 🤔:</p>\n<pre><code>k = 3\nsample_data = {\"a\": [\"1\",\"2\",\"3\"], \"b\": [] }\n\naverage_precision(data_true=sample_data, data_predicted=sample_data, k)\n</code></pre>\n<p>Perhaps you need to replace this…</p>\n<pre><code>if n_items_true == 0 or n_items_predicted == 0:\n    continue\n</code></pre>\n<p>with this?</p>\n<pre><code>if n_items_true == 0:\n    del data_true[key]\n    continue\nif n_items_predicted == 0:\n    continue\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1698820,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "02/20/2022 17:35:28",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a>, I think your modification is right, because customers having no purchases during test period are excluded from scoring. Hence, they should be excluded here to remove the effect on denominator <code>len(data_true)</code>. And, I think this issue is similar to what I comment in <a href=\"https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12\" target=\"_blank\">this kernel</a>.</p>\n<p>If there's any misunderstanding, please correct me. Thanks a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689952,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "02/14/2022 15:52:09",
      "content": "<p>for those that prefer pandas/apply, here's an exact implementation of the <a href=\"https://www.kaggle.com/nroman\" target=\"_blank\">@nroman</a>'s </p>\n<pre><code>def pandas_average_precision(\n    data_true: pd.Series,\n    data_predicted: pd.Series,\n    k: int,\n    norm: bool = True, \n) -&gt; float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if len(data_true)==0:\n        raise ValueError('data_true is empty')\n\n    # convert to df so we can use `apply`\n    eval_df = pd.DataFrame({\"true\": data_true, \"predicted\": data_predicted})\n\n    # replace na predictions with empty list\n    eval_df[\"predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: x if isinstance(x, list) else []\n    )\n\n    # getting the counts of true/predicted\n    eval_df[\"n_items_true\"] = eval_df[\"true\"].apply(len)\n    eval_df[\"n_items_predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: min(len(x), k)\n    )\n\n    # ignore zero true or predicted\n    non_zero_filter = (eval_df[\"n_items_true\"] &gt; 0) &amp; (eval_df[\"n_items_predicted\"] &gt; 0)\n    eval_df = eval_df[non_zero_filter].copy()\n\n    def row_precision(items_true, items_predicted, n_items_predicted, n_items_true, norm=norm):\n        n_correct_items = 0\n        precision = 0.0\n\n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            return precision / min(n_items_true, k)\n        else:\n            return precision / k\n\n    eval_df[\"row_precision\"] = eval_df.apply(\n        lambda x: row_precision(\n            x[\"true\"], x[\"predicted\"], x[\"n_items_predicted\"], x[\"n_items_true\"]\n        ), axis=1\n    )\n\n    return eval_df[\"row_precision\"].sum() / len(data_true)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1686637": "```\nimport numpy as np\n\n\ndef average_precision(data_true: dict, data_predicted: dict, k: int, norm: bool = True) -> float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if not data_true:\n        raise ValueError('data_true is empty')\n\n    average_precision_sum = 0.0\n\n    for key, items_true in data_true.items():\n        items_predicted = data_predicted.get(key, [])\n\n        n_items_true = len(items_true)\n        n_items_predicted = min(len(items_predicted), k)\n\n        if n_items_true == 0 or n_items_predicted == 0:\n            continue\n\n        n_correct_items = 0\n        precision = 0.0\n\n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            average_precision_sum += precision / min(n_items_true, k)\n        else:\n            average_precision_sum += precision / k\n\n    return average_precision_sum / len(data_true)\n```\n\nThis competition has Mean **Normalized** average precision score as a metric, so `norm` should be set to `True`",
    "1689676": "Thank you for sharing, @nroman! Did you do tests 😄?",
    "1689909": "nroman\nFor the following code, output is 0.5. 🤔:\n```\nk = 3\nsample_data = {\"a\": [\"1\",\"2\",\"3\"], \"b\": [] }\n\naverage_precision(data_true=sample_data, data_predicted=sample_data, k)\n```\n   \n    \n   \nPerhaps you need to replace this...\n```\nif n_items_true == 0 or n_items_predicted == 0:\n    continue\n```\nwith this?\n```\nif n_items_true == 0:\n    del data_true[key]\n    continue\nif n_items_predicted == 0:\n    continue\n```",
    "1689952": "for those that prefer pandas/apply, here's an exact implementation of the @nroman's \n\n```\ndef pandas_average_precision(\n    data_true: pd.Series,\n    data_predicted: pd.Series,\n    k: int,\n    norm: bool = True, \n) -> float:\n    \"\"\"\n\n    :param data_true: items that were actually purchased by user\n    :param data_predicted: items we recommended to user\n    :param k: number of candidates (predictions)\n    :param norm: nomalization flag\n    \"\"\"\n    if len(data_true)==0:\n        raise ValueError('data_true is empty')\n\n    # convert to df so we can use `apply`\n    eval_df = pd.DataFrame({\"true\": data_true, \"predicted\": data_predicted})\n    \n    # replace na predictions with empty list\n    eval_df[\"predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: x if isinstance(x, list) else []\n    )\n    \n    # getting the counts of true/predicted\n    eval_df[\"n_items_true\"] = eval_df[\"true\"].apply(len)\n    eval_df[\"n_items_predicted\"] = eval_df[\"predicted\"].apply(\n        lambda x: min(len(x), k)\n    )\n    \n    # ignore zero true or predicted\n    non_zero_filter = (eval_df[\"n_items_true\"] > 0) & (eval_df[\"n_items_predicted\"] > 0)\n    eval_df = eval_df[non_zero_filter].copy()\n    \n    def row_precision(items_true, items_predicted, n_items_predicted, n_items_true, norm=norm):\n        n_correct_items = 0\n        precision = 0.0\n        \n        for item_idx in range(n_items_predicted):\n            if items_predicted[item_idx] in items_true:\n                n_correct_items += 1\n                precision += n_correct_items / (item_idx + 1)\n\n        if norm:\n            return precision / min(n_items_true, k)\n        else:\n            return precision / k\n            \n    eval_df[\"row_precision\"] = eval_df.apply(\n        lambda x: row_precision(\n            x[\"true\"], x[\"predicted\"], x[\"n_items_predicted\"], x[\"n_items_true\"]\n        ), axis=1\n    )\n    \n    return eval_df[\"row_precision\"].sum() / len(data_true)\n```",
    "1698820": "Hi @jacob34, I think your modification is right, because customers having no purchases during test period are excluded from scoring. Hence, they should be excluded here to remove the effect on denominator `len(data_true)`. And, I think this issue is similar to what I comment in [this kernel](https://www.kaggle.com/kaerunantoka/h-m-how-to-calculate-map-12).\n\nIf there's any misunderstanding, please correct me. Thanks a lot."
  },
  "source": "meta"
}